A production Gen-AI video pipeline, built mid-project
Case Study
Stood up a production AI video pipeline from zero during a live client project: tooling, deployment, and artist training.
- The problem
- A national 30-second spot needed a brand's woodcut logo character animated in its own crosshatched style. By hand that would have cost five to ten times the budget [verify], and every revision would have taken a week or more. The studio had no AI video pipeline, and client material could not leave the building.
- What I did
- I ran a small pilot with two interns to find out how much of the job generative video could carry, then built the pipeline while the job was in production. I deployed it to eight workstations, trained a core team of six, and kept our 2D animators on hand for whatever it couldn't do.
- The outcome
- The spot delivered. Revisions that would have taken a week came back the same day, including character redesigns the client asked for in the final week. The pipeline has since run on at least four client projects and more than ten pitches.
Context
In mid-2025 the studio took on a 30-second national broadcast spot for a heritage food brand. The brief was to bring the brand’s logo character to life in the style it is printed in: woodcut, heavy crosshatching, every line looking carved.
Drawing that by hand was the obvious route, and it didn’t fit. My estimate put it at five to ten times the budget [verify basis]. The bigger problem was revisions. A change to a hand-animated shot in that style is a week of work at minimum, and commercial projects never really lock. Agency notes and client notes keep arriving until the day of delivery.
I had done a lot of work with image models by then, but the studio had no local video capability. I had tested the hosted services and the results were limited. The owner made the call to let me try, on the understanding that we would get as far as we could with AI and finish the rest by hand.
Two constraints shaped everything after that. The pipeline had to be built while the job was running, and it had to run entirely on our own hardware so client material stayed in-house. [link: the private IP and public models post]
Approach
Before committing the job, I wanted to know how much of it the technology could carry. I spent the first stretch running pilot shots with two interns. Once those showed that a usable share of each shot could be generated, we committed, with the 2D and VFX teams as the fallback for the gaps.
I set the bar for “done” in layers instead of one finish line:
- Other artists running it within one to two weeks
- Shots delivered within one to two weeks of starting
- Revisions the same day, with no re-animation
- The spot delivered
I looked at hosted video services, full hand animation, and a puppet-style 2D hybrid. Hosted services meant sending client material out, and they capped resolution and control. Local open models gave us pose and depth inputs, custom resolutions, and settings we could tune, so that is where we went.
For the technically curious: Wan VACE for video, with still frames restyled first through SDXL with custom LoRAs and later through Qwen-Image.
Implementation
The work ran on two tracks. On the look side, I trained a custom still-image model on the brand style and had styleframes and character designs in front of the client within two to three days. On the shot side, we cut a 2D animatic, shot reference footage of live talent against it, and made a rough cut. For each shot we restyled the first and last frames to the approved look, generated the video between them guided by the footage, and passed the result to 2D for cleanup before the online edit.
Getting it onto other machines
This was most of the work. I installed the pilot workstations by hand with locked versions, then moved all models to a shared network library where new versions get new folders and nothing is ever overwritten. An update could never break a shot someone was halfway through. It is a habit I carried over from years of loading 3ds Max plugins off the network.
Workflows followed shared conventions so any shot opened on any of the eight workstations without relinking, and I pushed updates with PowerShell. We also moved our A6000 cards into the machines that needed them.
The interns and I built the workflows. Artists ran them and changed them only when a shot called for it. Output went through the same review as everything else at the studio: dailies every day, approval in our project tracking, then on to composite and edit.
Training
I ran a one-hour workshop, recorded it for people joining later, and wrote a syllabus to refer back to. After each artist had a day or two of hands-on time, I sat with them for an hour or two and explained how the models work underneath. The aim was that when something went wrong they could reason about why, instead of deciding the tool was useless.
Some did decide that at first. Enthusiasm didn’t track seniority either: some veterans took to it right away and some newer artists resisted. [link: how to get a creative team to use AI]
The interns had no production quota. They experimented full time and reported to me daily, and I folded what worked into the pipeline.
What broke
As shots got longer and resolution went up, a single generation went from minutes to somewhere between twenty minutes and an hour. That is too slow to iterate with. I brought in tiled diffusion, built upscaling passes, and cropped each shot to the smallest region that needed processing. None of that would have been available on a hosted service.
Halfway through, the client asked to push the style further than originally scoped. With hand animation I would have had to push back. Here we adjusted the look and kept going.
Results
The spot delivered [on schedule: verify]. The first full version went to the client less than two weeks after the shoot, and that included a lot of conventional post alongside the generated work.
The result I would point to first is the last week. The client changed character designs days before delivery, and we turned those changes the same day. By hand, the answer would have been no.
The generated shots were not final. On average they got about 85% of the way there [verify], and two 2D animators spent roughly a week and a half bringing the character details up to brand standard. Fully hand-animated, the same spot would have been [number] animator-weeks.
Six people were proficient by the end of the project, four artists and two interns, and three more learned it afterwards.
Since then the pipeline has run on at least four client projects, in some cases replacing a reshoot [one concrete example]. It has been part of more than ten pitches, at least five of them won. A few AI motion tests made for pitch decks led clients to commission extra social content, which we then produced in traditional 3D. The same tools now get used for look development, textures, and footage cleanup.
Lessons
Every model we used on that job has since been replaced. What lasted was the process: model storage that never overwrites, shared workflow conventions, and a team that understands the tools well enough to pick them up without being asked.
What I would repeat: pilot before committing, keep a manual fallback alive, give a couple of people room to experiment with no delivery pressure, and teach how it works instead of which buttons to press.
What I would change: build the custom tooling sooner. The versioning, network file handling, and project-aware naming I later released as comfyui-pipedream would have saved time on the first job. I would also have artists lead workshops within the first year. People hear it differently from a peer.
When the spot aired, some viewers objected to the use of AI. We published a behind-the-scenes piece showing how the spot was assembled and how much of it was artists’ work, and the agency shared it on their own channels. [Your view on disclosure: what you would do up front next time.]