Work / Conductor
One product, 75+ generative models: the integration layer for an AI video platform
More than 75 generative model configurations from three vendors behind one product surface, with every generation's cost recorded and billed to match the vendor.
Challenge
Creators write a prompt, attach reference images, video or audio, and generate with many third-party video and image models side by side. Then they cut the results together on a shared timeline. New models ship every few weeks from several vendors, and each has its own inputs, limits, pricing and failure modes. The product had to absorb that churn without a UI rewrite for every model, and without a bill that drifted from the quote.
Approach
01Models are data, not code
Every model is a declarative record: vendor, variant, accepted inputs with their limits, and a pricing matrix. One generic form renderer turns those records into the UI, and the same records drive validation on both client and server. Adding a model is a data change delivered by a sync command. It needs no per-model UI work and no database migration, and a bad input fails before it reaches a paid vendor call.
02Route by capability, not by vendor
Variants of one model family (text-to-video, image-to-video, reference-to-video, edit) sit behind one tile. The UI asks the schemas which inputs each variant accepts. When more than one vendor sells the same family, the user picks the API inside the tile. The list grew to more than 75 model configurations across three vendors plus an in-house diffusion pipeline, and users still see one tile per model family.
03Treat vendor APIs as slow and unreliable by default
Each model gets a wall-clock polling budget, and transient network errors are retried inside it. Outputs of several gigabytes stream to disk in 1 MB chunks instead of loading into memory. Jobs are acknowledged late, so a worker that crashes hands its job to another. An hourly sweep fails anything abandoned, and users are charged only on success. A vendor slowdown or a worker crash no longer leaves a generation stuck in "Creating".
04Bill what the vendor bills
Users see a cost estimate before they submit. Each finished generation records its own cost at sub-cent precision, rounded once. Where the vendor reports the true cost, that figure is billed, and rates were re-derived from measured vendor charges. Costs roll up per folder and per project, and failed jobs show "No charge". Because the vendors offer no spend API, stored per-generation cost also gives the business a per-model, per-month view of inference spend.
05Make concurrent timeline editing converge
The timeline is stored as OpenTimelineIO data inside a CRDT document held by a sync service. (A CRDT is a data type that merges concurrent edits without a central lock.) Every clip carries a stable ID, so a drag-and-drop move keeps its identity. After each merge, a normalization pass removes duplicate clips and resizes gaps. It only deletes or resizes, never inserts, so any two clients that repair the same state reach the same result. Characterization tests of real merge conflicts back the rules.
06Ship with coding agents under hard gates
One senior engineer led architecture and review, working with a frontend developer from the Eastern European team. Coding agents did much of the implementation. The agents read a single checked-in contract and architecture guide, and repeatable skills run every lint, type, test and codegen gate before a pull request opens. An AI reviewer is limited to correctness, security, performance and architecture. CI is the only merge gate, so every change passes the same checks whether an agent or a person wrote it.
Results
- More than 75 generative model configurations from three vendors behind one product surface, plus an in-house diffusion pipeline.
- Adding a model is a data change with no per-model UI work or database migration.
- Vendor slowdowns and worker crashes no longer strand generations. Users are charged only on success.
- Charges come from each generation's stored cost, billed at the vendor-reported figure where the vendor provides one.
- A project-content load that ran more than 3,000 database queries (about 25 s) was fixed by removing per-item work from list reads. Query-count tests show a 20-item media listing falling from 107 queries to 9, and a 230-item timeline listing from 237 to 7.
- CDN URL signing dropped from about 200 ms to about 4 ms per URL. The slow signing had been causing health-check restarts.
- 75 tagged releases in about 10.5 months.
Capabilities
AI Engineering
- Integration and routing across many models from several generative video and image vendors
- Inference cost control: estimates before submit, per-generation cost records, billing at vendor-reported cost
- Reliability engineering for long-running third-party inference jobs
- AI-assisted development: coding-agent contracts, gated agent skills, a scoped AI reviewer
Software Engineering
- Architecture and API design: declarative model schemas, a GraphQL API with batched loaders
- Real-time collaboration: CRDT timeline sync with permission checks on every message and live access revocation
- Performance: removed N+1 query patterns, added connection pooling and a signed-URL cache
- Security and access control: role-based permissions, cross-project access fixes, secret scanning, log redaction
- Cloud infrastructure and IaC on AWS: container services, managed Postgres and Redis, a signed-URL CDN, zero-downtime deploys
- CI/CD: sharded browser and unit test suites, parallel deploys
- Data pipelines: two-way cloud-storage sync with loop prevention and layered deduplication