Work / TenvosAI

Stop paying for idle GPUs, and rebuild the platform so each customer is isolated

GPU hosts that shut themselves off with an independent backstop, idempotent inference, and a platform in Terraform that can be parked with one variable.

Client
TenvosAI
Context
A voice-AI startup that detects fatigue and impairment from speech.
Industry
Voice AI
Role
Fractional CTO
Service lines
AI Engineering, Software Engineering

Challenge

The training job, inference service, API, web apps and database all ran on one server per environment, and a nightly GPU training host could stay on for many hours after its job finished. The company was adding its first external API customer and needed tenant isolation, reliable deploys and predictable GPU spend.

Approach

  1. 01A training host that turns itself off, with a second line of defense

    A scheduler starts the GPU host each night, the host stops itself when the job ends, and the job has a hard time cap. A separate function checks every 15 minutes and force-stops any host still running after 90 minutes. One failure mode is not enough: in the history, a missing permission kept a box up for about 21 hours, and a boot failure stranded another for 13.

  2. 02Less work per training run

    Data that was fetched once per sub-model is now fetched once per tenant, and a shared model is built once and reused without a GPU. The cheapest GPU minute is the one you don't schedule.

  3. 03Idempotent inference requests

    Each recording now runs through the GPU once. Client retries replay the stored result instead of starting a second 5–50 second pass. Mobile clients retry, so without this the bill scales with bad network conditions.

  4. 04GPU inference on its own capacity pool, deployed without an outage

    The inference service moved to a dedicated GPU pool that temporarily adds a second host during deploys. Model files are baked into the image, so a new host does not fetch them in the middle of a request. Those cold-start fetches had been pushing requests past the time limit.

  5. 05Infrastructure as code with a single "park" switch

    The whole platform is defined in Terraform. One variable scales every service and host pool to zero and turns off the nightly schedule, while the runaway-host safety check stays on. An early-stage company can pause spend in one reviewed change and restore it the same way.

  6. 06Deploys that prove they landed

    Each deploy pins the image to the commit, and the pipeline confirms the service is running the new revision rather than trusting a green status. An automatic rollback can look like a success unless you check.

Results

  • The nightly GPU training host stops when its job ends, and an independent check force-stops any host still running after 90 minutes.
  • Training runs fetch data once per tenant instead of once per sub-model.
  • Client retries replay a stored result instead of paying for a second GPU pass.
  • Every deploy is pinned to a commit and confirmed against the running revision.
  • Automated test coverage grew from no test pipeline to 463 tests on the core API.

Capabilities

AI Engineering

  • GPU and inference cost optimization
  • ML training and serving infrastructure
  • AI-assisted development (coding-agent workflows)

Software Engineering

  • Cloud infrastructure and IaC
  • CI/CD and DevOps
  • Architecture and API design
  • Security and access control
  • Reliability, observability and incident debugging
  • Technical leadership

Stack

  • AWS (ECS on EC2, GPU instances, RDS for PostgreSQL, S3, KMS, Lambda, EventBridge, ALB, WAF, CloudWatch)
  • Terraform
  • Docker
  • GitHub Actions
  • PostgreSQL
  • PgBouncer
  • Python
  • Django
  • React
  • React Native