Skip to content
Ease AI
← All insights

What is your AI actually doing? The reporting layer most builds skip

The moment a model answers on your behalf, you need to know what it said, how often it refused, and whether the bill was reasonable.

August 25, 20266 min read

The moment an AI model answers a question on your behalf, you own what it said. You also own what it cost, who used it most, and whether it refused to answer when it should have escalated instead.

Most builds ship without any way to answer those questions.

The gap between deployment and accountability

You can buy a chatbot that answers customer questions, automates intake forms, or drafts replies to support tickets. The vendor will tell you it works. They may show you a demo. They will not, by default, give you a log of every prompt it received, every response it generated, or a breakdown of which users cost you the most money.

That gap matters less when the AI is internal. If your sales team uses a tool to draft follow-up emails, you can watch the drafts before they go out. You can ask the team whether it helps. You can turn it off.

It matters considerably more when the AI is customer-facing. A chatbot on your website does not send you a draft. It replies. Your customer reads it. If it gets the price wrong, invents a policy, or refuses to answer a question a human could have handled, you will not know unless the customer complains. Most will not. They will leave.

What a reporting layer actually tracks

A proper reporting layer logs every model call. Each log entry captures the prompt, the response, the model used, the token count, the latency, and the cost. It attributes that cost to a specific user and a specific feature. It tracks refusal rates—how often the model declined to answer. It tracks escalation rates—how often it handed off to a human. It enforces spend caps and sends alerts when usage exceeds thresholds you set.

It also produces an exportable audit trail. That matters when a customer asks what the system said to them three months ago, or when you need to demonstrate compliance with a privacy regulation, or when your accountant wants to reconcile your AI bill against actual usage.

None of this is automatic. If your build does not include it, you are flying blind.

Per-user and per-feature cost attribution

AI models charge by the token. A single customer conversation might cost two cents. A bulk document review might cost twelve dollars. If you are running both through the same system, you need to know which feature is driving the bill.

Per-user tracking tells you whether one customer is responsible for half your monthly spend. That might be fine—they might be your largest account. Or it might mean someone is testing your system to see how much they can extract before you notice.

Per-feature tracking tells you whether the chatbot you built to save time is actually cheaper than the human process it replaced. If it is not, you have a decision to make. The data lets you make it with your eyes open.

Refusal and escalation rates

A model refuses when it cannot answer a question. It escalates when it hands off to a human. Both are normal. The rate at which they happen is not.

If your chatbot refuses to answer fifteen percent of questions, you need to know. That might mean the training data is incomplete, the prompt is too restrictive, or the questions are outside the scope you planned for. All three are fixable, but only if you know the refusal rate.

If your chatbot escalates forty percent of conversations, it is not saving time. It is adding a step. You might still want it—it could be triaging effectively, routing complex questions to the right person faster than a form would. But you should know the number before you decide whether the build was worth it.

Spend caps and alerts

AI costs scale with usage. That is manageable when usage is predictable. It is not manageable when a misconfigured feature calls the model in a loop, or a new user runs a bulk job that costs three hundred dollars in an afternoon.

Spend caps prevent runaway costs. You set a monthly or per-user limit. When usage hits the cap, the system stops calling the model. It logs the event. It alerts you. You decide whether to raise the cap or investigate why usage spiked.

Alerts let you respond before the cap is hit. You set thresholds—fifty percent of monthly budget, eighty percent, ninety percent. The system emails you. You check the logs. You see which user or feature is responsible. You adjust.

Without caps and alerts, you find out when the bill arrives. By then, the money is spent.

The audit trail a customer eventually asks for

A customer will eventually ask what your system said to them. They might phrase it politely. They might phrase it through a lawyer. Either way, you need an answer.

The audit trail is a timestamped, exportable record of every interaction. It includes the prompt, the response, the model version, and the user identifier. It is stored separately from the application database, so it survives even if the application is rebuilt or replaced.

This is not paranoia. It is basic record-keeping. If you are shipping AI to your own customers, you are making statements on their behalf. You need to know what those statements were.

When not to build any of this

Most small businesses do not need a custom AI build at all. If you are asking whether AI could help your team draft emails faster, the answer is probably yes, and the tool is probably ChatGPT or something similar. You do not need logging. You do not need cost attribution. You do not need a reporting layer. You need a subscription and an afternoon of training.

This matters when you are shipping AI to your own customers. If the AI is customer-facing, if it answers questions or completes transactions on your behalf, if it runs without supervision, then you need to know what it is doing. That is when the reporting layer stops being optional.

If you are not shipping to customers, you probably do not need it. Do not pay for infrastructure you will not use.

What we build when you do need it

When a build does require a reporting layer—because it is customer-facing, because cost attribution matters, because you need an audit trail—we include it as part of the custom AI application build. Every model call is logged with prompt, response, model, tokens, latency and cost. Cost is attributed per user and per feature. Refusal and escalation rates are tracked. Spend caps and alerts are configured. The audit trail is exportable.

That sits in the $20,000-$35,000 range, delivered over eight to twelve weeks. It is not cheap. It is appropriate when the AI is doing something that matters enough to need accountability.

If you are based in Ontario and the build qualifies as experimental development—genuine technological uncertainty, not routine implementation of an existing platform—the enhanced SR&ED credit may apply. The refundable expenditure limit was raised from $3 million to $6 million for taxation years beginning on or after 16 December 2024, and capital expenditures including development-related software and cloud infrastructure were reinstated. That is a tax credit, not a grant, and it requires documentation. It does not apply to most projects. When it does, it is worth claiming.

You can see the full breakdown on the pricing page, or reach me directly at info@easeaiworks.com.

This article was drafted with AI assistance against a research brief and published automatically. Every figure links to a primary source. If you find an error in it, tell me and I will correct it — that offer is the point.

While you are here

Two minutes to find out where you actually stand.

Eight questions, a score, and a likely cost range — shown before anyone asks for your email address.

Take the scorecard

Got a question about any of this?

Ask it. I reply within one business day, and I do not put you into a sales sequence for asking.

Replies within one business day, from me directly.