AhmsDevAI Cost

Example report · example, made-up company

AI cost audit example report

This is a complete AI cost audit for a made-up company, built from made-up usage data with the same method a real audit uses. It is not a customer. Provider and model names are replaced with letters.

Last updated:

AI cost audit: Example Co.Example, made-up company

1. Summary

Example Co. spends $6,400 a month on AI APIs across 2 providers and 4 features. I found 5 changes. Together they are worth an estimated $1,395 to $2,885 a month.

Start with number 1. It is the largest and the easiest to test. Numbers 3, 4 and 5 are small code changes with low risk.

2. Where the money goes

ProviderPer monthShare
Provider A$4,80075%
Provider B$1,60025%
Example, made-up company
Feature (from request tags)Per monthShare
Support chat assistant$2,60041%
Ticket tagging$1,50023%
Document summaries$1,30020%
Internal search$6009%
Untagged$4006%
Example, made-up company

6% of spend had no feature tag. Tagging it would make the next review sharper.

3. Ranked changes

#ChangeEstimated per monthEffort
1Move ticket tagging to a smaller model
Ticket tagging, Provider A
$640 to $1,275Low
2Trim boilerplate from summary inputs
Document summaries, Provider A
$260 to $520Medium
3Re-embed only changed documents
Internal search, Provider B
$200 to $400Low
4Cache the repeated system prompt
Support chat assistant, Provider A
$195 to $390Low
5Cap retries on timeouts
All features, Provider A
$100 to $300Low
Example, made-up company

Total: $1,395 to $2,885 a month (estimated).

Want this for your own AI spend?

Request my audit

4. Method notes

  1. Move ticket tagging to a smaller model

    Tagging costs $1,500 a month on a frontier model. A smaller model costs about 15% as much for short classification tasks, so a full switch would cut about $1,275. The low end assumes only half the volume can move. How to check: Run 200 labeled tickets through both models and compare tags.

  2. Trim boilerplate from summary inputs

    In a sample of 50 calls, 20% to 40% of each input was headers, signatures and repeated legal text. Input is most of this feature's cost, so the range follows that share of $1,300. How to check: Compare 30 summaries before and after trimming.

  3. Re-embed only changed documents

    The index re-embeds every document each night. About two thirds of documents did not change between runs. The low end allows for documents that change often. How to check: Confirm search results match on 20 common queries.

  4. Cache the repeated system prompt

    The same long system prompt goes out with every chat turn and makes up about 30% of this feature's cost. Cached input is billed at about half price on this provider. How to check: None expected. Output does not change.

  5. Cap retries on timeouts

    From March 14 to 16, timeouts triggered up to 6 retries per call and added about $300. Capping retries at 2 stops a repeat. The low end assumes spikes like this happen every few months. How to check: Watch error rates for a week after the change.

5. What this audit did not check

Two months of data only, so seasonal swings may be missed. Quality was not tested; each change lists the test to run first. Prices are the providers' list prices on the date of the report.

6. Next

A 5 to 8 minute video walks through this report. Send questions in writing within 14 days and I answer them in one written round.

Back to home