AI infrastructure & unit economics · independent consulting
Nobody owns a cost they can’t see.
I find the AI and cloud spend nobody owns, attribute it to the teams that can act on it, and rank the waste by dollars. I worked on cloud and AI infrastructure in big tech, at Apple, Meta and Roku, and I build unalloc, an open-source tool for exactly this.
Questions welcome · ready to hire? rates & fit check
Tim Urista · Connect on LinkedIn
AWS spend attributed
$10–12M/quarter
The cost-attribution pipeline I built at Roku, mid-way through a Kubernetes migration.
Who pays for the KV cache?
arXiv:2609.24991
My preprint on splitting shared GPU bills, checked on a rented H100. Public and citable, not peer reviewed.
Production AI on my own invoice
Since Feb 2024
TrendVesting, where one cheap measurement overturned 2.5 years of beliefs about its confidence scores.
Work you can inspect
Built, run and measured by me
Products and tools I operate myself. Each one links to the write-up, with the commits behind it.
Open source · unalloc
Find the AI spend nobody owns
Joins Kubernetes allocations with OpenAI, Anthropic and gateway bills, ranks the unowned rows by dollars, and fails CI when too much spend has no owner.
Read the case →
SaaS I run · TendForm
A HIPAA-capable product, shipped solo
A form builder with an MCP server and a HIPAA tier, priced against its own cost stack and checked live before launch.
Read the case →Production AI · TrendVesting
Measure before you believe the model
Two and a half years of signals, then one calibration check: the win rate barely moved at any confidence bar. The scores meant nothing until they were measured.
Read the case →What I do
Spend less on AI. Get more out of it.
Three kinds of work, one discipline: measure the unit that matters, then change the system until the number moves.
01 · Attribute
Find out what your AI actually costs
Per request, per customer, per accepted output. Map spend to owners so someone can act on it, then rank the waste by dollars.
Learn more →02 · Manage
Ship faster with AI without shipping worse
Scoped work, review gates, tests, and verified-live checks that catch what fast delivery breaks. The same process I run on my own products.
Learn more →03 · Rebuild lean
Replace the expensive thing with a smaller one
Websites and internal tools rebuilt AI-native: less code, lower run cost, and a deploy pipeline you own.
Learn more →Writing · Human in the Loop
What the commit history actually says
Field notes on managing AI-assisted software delivery: how I scope the work, what I check before anything ships, and the bugs that got through anyway.
- Guide
Sep 15, 2026 · 11 min · unalloc
Who pays for the KV cache? Designing showback for shared LLM inference
A guide for platform and FinOps leads who have to split shared AI spend: which meter to charge by, where idle capacity goes, how to give every dollar one path into the ledger, and what a joined ledger can answer. Backed by measurements from unalloc, my open-source cost-attribution tool.
- Part 10
Sep 14, 2026 · 11 min · unalloc
unalloc: checking a cost-attribution tool before anyone quotes its numbers
How I scoped and checked unalloc, an open-source tool that joins Kubernetes and LLM provider bills: nine defects caught before release, a budget-capped H100 run with a verified teardown, a paper build that refuses silent typos, and a headline I softened because no meter is ground truth.
- Part 2
Sep 13, 2026 · 11 min · TrendVesting
TrendVesting: validating 2.5 years of production AI
How I review and validate an AI trading system I own: change records with production numbers, invariant tests, reconciliation audits and calibration checks, and the numbers they overturned.
Free lesson · 10 minutes · no sign-up
Ready, warm, useful
Why a replica can pass its readiness check while the first users still wait. An interactive timeline separates readiness, a cold prefix cache and the queue.
Try it
AI cost teardown
Describe a workload in a sentence. Get a monthly estimate, a unit cost, and the levers, the way I’d sketch it on a first call.
Estimates from a language model. Useful for orientation, not a quote.
AI Cost Teardown
Describe an AI or cloud workload in plain English. You'll get a fully-loaded monthly estimate, a per-unit cost, and the highest-leverage ways to cut it. Live, powered by Claude.
Rates
Published prices. Cash-pay.
Invoiced directly, paid by ACH or card. No platforms or agencies in the middle.
Advisory hours
$250/hr
2-hour minimum
Executive Cost Briefing
$2,500
90 minutes
Unit Economics Sprint
from $5,000
1 day
AI Cost & Unit Economics Audit
$12k–$25k
1–2 weeks
ML Cost & Reliability Evals
$15k–$30k
2 weeks
AI-native rebuild
Fixed quote
After a Sprint