Studios working on unreleased shows often can't send scene files or pipeline code to cloud AI tools. Client security terms rule it out. Houdini, the most technical tool in the pipeline, is left without AI help.
I founded Nilo AI to fix that. The agent runs inside a live Houdini session on the studio's own hardware. It is a native desktop app (Tauri/Rust) with a local model backend, connected to Houdini through the Model Context Protocol (MCP), a standard way for AI models to call real tools. It has 188 of them across 23 categories. It builds and repairs node networks, writes VEX and checks it against the running session, sets up Pyro, FLIP, RBD and Vellum simulations, and generates character motion as editable APEX animation layers.
The model learns from a pipeline that turns real Houdini scene files into training examples. Each example uses the same tool format as the app, so the training maps directly to what the agent does in production. A corruption engine breaks working node graphs on purpose, which gives repair pairs with known correct answers. Only a third of those examples name the fault. The rest describe a symptom or give no hint at all, which trains the model to find the problem itself.
The production model is Qwen3.8-27B, an open-weight model, fine-tuned with QLoRA. QLoRA trains a small set of adapter weights on top of a compressed copy of the model, so the 27 billion base parameters stay untouched. One pass over 9,500 examples (17.2M tokens) took 7.1 hours on a single A100 GPU and cost $34. At 4-bit, the finished model fits on one workstation GPU.
Six models were tested under the same conditions on 2,221 held-out examples from sources the training data never used. Every model saw three worked examples first, which helps the untuned models the most. Every generated VEX snippet was compiled and run in a live Houdini 22 session. The models were compared on the same examples with McNemar's exact test, with Holm correction for the ten comparisons.
The tuned model compiled 89.6% of its VEX. That is more than any other model tested and above the 83.3% of the reference answers. On graph repair (72.0%) and on finding faults with no hints (54.5%), it showed no statistically detectable difference from Claude Opus 5 at this sample size. It beat every untuned open-weight model tested. The same pipeline, applied unchanged to two generations of Qwen, raised graph repair from 14% to 72%.
Claude Opus 5 and Sonnet 5 are the frontier reference in every benchmark. Claude as a judge for open-ended Houdini answers, which automatic scoring can't grade, is in progress and runs in the next training stage.
The next experiment is the most important one: testing whether the same gains hold on a studio's own scenes and conventions.
Planned: an opt-in Claude tier for studios with looser data rules, next to the fully on-premise default.
Below are the product overview and the full technical report, slide by slide. To see the report in full quality, you can download it by clicking here.