Work / Optimizing Mind
From a compiled model to a two-environment production service, plus a site built for search and AI crawlers
A proprietary model behind a private, scale-to-zero service, released through CI to staging and production built from the same Terraform.
Challenge
The company's model shipped as a proprietary compiled binary. Customers needed a secure web API in front of it, and the team needed a repeatable way to test and release changes. The marketing site also needed to load fast, get its videos indexed, and be readable by AI assistants.
Approach
01Separate the model from the public API
The model runs in its own container service with no public ingress. It scales to zero when idle and scales out on concurrent requests. Model compute costs nothing between training runs, and the model is never exposed to the internet.
02Ship the model through CI, not by hand
A two-stage Docker build packages the model with a small Go server. Every push builds an image tagged with the commit and deploys a new revision to the matching environment, so every running revision traces back to one commit.
03Build staging and production from the same Terraform
Both environments get a private network, PostgreSQL on a private subnet, Key Vault secrets read by managed identity, and centralized logs. Production matches staging by construction, and no secret values sit in container configuration.
04Make configuration fail loudly
Terraform passed "true" where the app expected "True", so email silently stopped in both environments. The fix was a strict parser that rejects unknown values, plus a test. Bad configuration now stops the deploy instead of hiding in production.
05Give each video its own watch page and a video sitemap
A "noindex" header added to fix a Search Console warning blocked video indexing. It was reverted within three days and replaced with the documented video-sitemap mechanism. Search engines need a page per video, and header rules on video files can hide the video from indexing.
06Publish an llms.txt and open the site to AI crawlers
The site has a curated plain-text summary, schema.org markup and explicit crawler rules. A rule in the repo keeps the summary in sync with page content. AI assistants cite what they can parse, and they trust it more when it matches the page.
Results
- Staging and production built from the same Terraform, with every running revision traced to a commit.
- Versioned releases with guard checks.
- The model service scales to zero when idle and has no public ingress.
- Demo video cut from 15 MB to 2.7 MB. Explainer video cut from 219 MB to 40 MB.
- Mobile poster image cut from 59 KiB to 23 KiB. Text contrast raised to WCAG 4.5:1.
- Each video has its own watch page and is listed in a video sitemap.
- Every lead form submission is saved to the database before email is sent.
Capabilities
AI Engineering
- MLOps: packaging a proprietary model into a container and shipping it through CI
- ML serving infrastructure: internal-only model service, scale to zero, health probes
- LLM discoverability: llms.txt, structured data and AI crawler access
- AI-assisted development: repo instructions for coding agents, agent-authored commits
Software Engineering
- Cloud infrastructure and infrastructure as code across two environments
- CI/CD with per-commit images and branch-based deploys
- Release management: version bumps, release pull requests, guarded tagging
- Reliability: tracing configuration drift between infrastructure and application
- Security: private networking, managed identity, secrets in a vault
- Web performance, technical SEO and accessibility
- Lead capture that saves every submission before sending email