A practical architecture for generating report narratives with AI — and the guardrails that keep it from turning a decline into a success story.
By David Henderson • Unwired Web Solutions • Updated September 2026 • 11-minute read
|
DIRECT ANSWER To automate a monthly report narrative safely, keep the language model inside a controlled pipeline: calculate metric changes in code, attach operational context, require every claim to quote a supplied value, validate the draft against the source data, and route anomalies to a human. The LLM writes; it does not decide what happened. |
|---|
The problem: the dashboard gets automated, but the explanation doesn’t
Reporting dashboards get automated far more often than the narrative that goes with them. The extraction and visualization work is usually already scripted — GA4, Search Console, and ad-platform data land in a dashboard automatically. The explanation of what it means almost never is.
Clients rarely need another chart. They need someone to say what changed, why it may have changed, what the team did about it, and what happens next. That’s exactly the kind of writing a language model is well suited to draft — and exactly the kind of writing where an ungoverned model will confidently say the wrong thing.
|
What is an automated report narrative? An automated report narrative is a written interpretation of structured performance data generated by software. A reliable system combines deterministic calculations, business context, constrained language generation, and a verification step. It should explain supplied evidence — not infer a convenient story from it. |
|---|
A production-grade architecture: Python, analytics data, and campaign context
The pattern below is the one I’d recommend for this: a scheduled job pulls current- and prior-period metrics, calculates percentage changes in code, adds completed and planned work from wherever your team tracks it (Jira, Asana, or similar), and sends a compact structured payload to a language model to draft from.
Recommended stack
| Layer | Implementation | Purpose |
|---|---|---|
| Runtime | Python | Extract, normalize, and validate metrics |
| Analytics | GA4 Data API + Search Console API | Sessions, engagement, key events, clicks, impressions, position |
| Operational context | Project-tracking tool (e.g. Jira) | Completed work and the next sprint |
| Language model | A current, supported model API | Draft the narrative from the supplied payload |
| Schedule | A scheduled job runner (e.g. GitHub Actions) | Run the workflow each reporting period |
Model usage cost for a run like this is typically small relative to the analyst time it replaces — comfortably in the range of cents per report, not dollars — though the exact figure depends on payload size and which model you run.
An illustrative shape for the payload:
def build_report_payload(current, prior, context):
return {
"metrics": {
"organic_sessions": metric(current["sessions"], prior["sessions"]),
"organic_conversions": metric(current["conversions"], prior["conversions"]),
},
"completed_work": context["tracked_completed_items"],
"planned_work": context["tracked_next_period_items"],
"data_quality_flags": context["tracking_anomalies"],
}
The important design choice is that the model doesn’t calculate the deltas. Code does. The payload supplies the current value, prior value, percentage change, data-quality flags, and work context as separate fields — which makes the draft easier to test and removes any ambiguity about whether a number rose or fell.
The failure mode this has to guard against
Here’s the risk made concrete with a hypothetical example. Imagine a monthly payload showing:
-
Organic sessions: 12,450 (+4.2%)
-
Organic conversions: 310 (−12.0%)
-
Form submissions: 82 (−22.5%)
An unconstrained model asked to draft an upbeat client update from this could easily lead with the traffic increase and describe performance as “solidifying steady conversion volume” — a sentence that’s fluent, positive, and directly contradicted by the conversion numbers in the same payload. The model wouldn’t have fabricated a number. It would have fabricated the meaning of the numbers — a false interpretation, not a false statistic.
|
THE OPERATIONAL RISK A reporting hallucination isn’t limited to a made-up metric. It can also be a false causal claim, an unsupported attribution, or positive framing that contradicts a negative KPI. In a client report, all four are credibility failures. |
|---|
This is why “ask the model to be accurate” is not a control. Language models produce plausible text from patterns. They don’t independently know which KPI a client considers decisive, whether tracking was stable, or whether a completed task caused a result. Those facts have to be encoded, tested, or withheld.
What stops an AI report from inventing a win?
A longer prompt by itself won’t fix this. The reliable fix is two layers of deterministic control plus a human release gate.
1. Validate direction before generation
A pre-processing step classifies every KPI before the payload reaches the model. When a core metric moves beyond an agreed threshold, the pipeline injects an explicit constraint. Thresholds should be account-specific — a five-percent change may be meaningful for one KPI and noise for another.
def generate_constraints(metric_deltas):
constraints = []
for kpi, change in metric_deltas.items():
if change < -5.0:
constraints.append(
f"CRITICAL: {kpi} decreased by {change:.1f}%. "
"State this decline in the first paragraph. "
"Do not describe it as positive, stable or optimal."
)
return "\n".join(constraints)
2. Constrain claims and attribution
The system instruction should require exact values, separate completed work from proven causes, and define a failure response for missing or anomalous data. Completed work can show that something happened; it can’t by itself prove that it caused a result.
| REPORTING RULES |
|---|
|
3. Reconcile the draft after generation
A post-generation validator should extract percentages and metric names from the draft, compare them with the source payload, and scan for prohibited positive framing near negative deltas. It should also check that each material decline appears in the narrative. A mismatch should fail the run rather than trigger an automatic rewrite loop that could quietly hide the error.
| Control | Pass condition | Failure action |
|---|---|---|
| Numeric reconciliation | Every stated number exists in the payload | Stop and review |
| Directional consistency | Language matches positive/negative direction | Stop and review |
| Material KPI coverage | Every threshold breach is disclosed | Stop and review |
| Attribution discipline | Causal claims have approved evidence | Remove claim or review |
| Data completeness | Required fields and tracking flags present | Use anomaly response |
4. Keep a human release gate
The final human review shouldn’t be ceremonial. Someone checks the headline conclusion, data-quality flags, causal language, and next actions before publication. The automation owns assembly and first-draft prose. The account team keeps accountability for interpretation and client communication.
What this approach buys you
Done well, a pipeline like this turns report narrative drafting from one of the more repetitive parts of the reporting cycle into a short review step. The actual time saved depends on account volume and how much of the current process is already structured — but the larger gain isn’t really about speed. It’s role clarity: staff stop spending the reporting cycle assembling sentences and spend it on exceptions, interpretation, and strategy instead.
A reusable framework for AI-generated performance commentary
-
Define the decision metrics. Identify the KPIs that change the client conversation and set materiality thresholds for each one.
-
Calculate outside the model. Normalize date ranges, percentages, and comparison logic in code — not in the prompt.
-
Separate facts from context. Store metrics, completed work, planned work, tracking anomalies, and approved causal notes as distinct fields.
-
Write a claim policy. Require exact values, plain treatment of declines, and explicit handling of missing data.
-
Validate the output. Reconcile numbers, direction, KPI coverage, and causal language against the payload.
-
Escalate instead of improvising. A failed validation should create a review task, not invite the model to keep guessing.
-
Log the evidence. Retain the input, model/version, prompt version, validator result, reviewer, and publication timestamp.
|
RULE OF THUMB Use an LLM for language transformation, not for arithmetic, source-of-truth selection, or final accountability. The more consequential the report, the more of the pipeline should remain deterministic. |
|---|
Model versions are dependencies
Anyone building this needs to treat the model itself as a versioned production dependency, not a fixed constant. As one concrete example of how fast this moves: Anthropic deprecated claude-3-5-sonnet-20241022 on August 13, 2025, and retired it on October 28, 2025. A pipeline built against a specific model ID needs a plan for what happens when that model is retired — pin the model ID explicitly, and rerun your evaluation set before deploying a replacement. A model swap can change tone, instruction-following, extraction behaviour, and cost even when the prompt is unchanged.
Keep a small test suite that includes a normal month, a strong decline, mixed metrics, missing data, a tracking anomaly, and an unsupported causal claim — and rerun it every time the underlying model changes.
Frequently asked questions
Can AI write a monthly marketing report narrative?
Yes. An LLM can turn structured performance data and campaign context into a readable first draft. For client-facing use, calculations should happen in code, claims should be constrained to supplied evidence, and the draft should pass automated validation plus human approval.
How do you prevent hallucinations in automated reports?
Don’t rely on prompting alone. Reconcile every number with the source payload, test whether the language matches the metric direction, require material declines to be disclosed, prohibit unsupported causal attribution, and stop the workflow when data is incomplete or anomalous.
Why not use a dashboard tool’s built-in narrative features?
A widget-level summary may not have access to historical account context, completed work, deployment notes, or custom data-quality flags. A custom pipeline can unify those sources and apply account-specific thresholds before language generation.
Can the model recommend what to do next month?
Yes, when planned tasks and decision rules are included as structured context. Recommendations should be labelled as planned actions — not presented as proven consequences of last month’s performance.
What happens when the GA4 Data API reaches a quota?
Use bounded retries with backoff, monitor property quota consumption, and fall back only to a clearly timestamped cache. Google’s Data API can return property quota details when requested. A stale fallback should be disclosed rather than silently treated as current data.
Does this eliminate the need for account managers?
No. It reduces repetitive drafting. Account managers still validate interpretation, explain uncertainty, discuss priorities, and own the client relationship.
Should FAQ schema be added to this article?
Not for the purpose of earning a Google FAQ rich result — Google deprecated that feature in 2026. Keep the FAQ because it answers real questions clearly; use Article and Person/Organization markup to describe the page and its author accurately.
The lesson: automate the draft, not the responsibility
Automating executive commentary isn’t about removing human accountability. It’s about moving the human from blank-page writer to reviewer, exception handler, and strategist.
A system like this can optimize for speed and fluent prose, or it can optimize for traceability — every important sentence pointing back to a metric, an approved context field, or a clearly labelled plan. Only the second version is safe to hand a client. That’s what makes the time savings usable without spending client trust.
For the broader operating philosophy behind this approach, read “There Are No 7 Prompts: 50 Things I Actually Do With AI Inside a Working Agency”. For the reporting foundation, see “The Client Dashboard That Replaced a Monthly Reporting Call”.
Sources and methodology
Product-status and platform claims in this article were checked against the following first-party documentation:
-
Google Analytics Data API overview
-
Google Analytics Data API limits and quotas
-
Anthropic model deprecations
-
Google: creating helpful, reliable, people-first content
-
Google: generative AI content guidance
-
Google Search documentation updates
About David Henderson
David Henderson is an SEO and digital marketing practitioner with more than 25 years of experience building, managing, and evaluating search, reporting, and content systems. His current work focuses on practical AI integration: where automation improves production, where it fails, and what still requires human judgment. He is the founder of Unwired Web Solutions and writes at davidhenderson.ca.