A practical architecture for generating report narratives with AI — and the guardrails that keep it from turning a decline into a success story.

By David Henderson • Unwired Web Solutions • Updated September 2026 • 11-minute read

DIRECT ANSWER

To automate a monthly report narrative safely, keep the language model inside a controlled pipeline: calculate metric changes in code, attach operational context, require every claim to quote a supplied value, validate the draft against the source data, and route anomalies to a human. The LLM writes; it does not decide what happened.

The problem: the dashboard gets automated, but the explanation doesn’t

Reporting dashboards get automated far more often than the narrative that goes with them. The extraction and visualization work is usually already scripted — GA4, Search Console, and ad-platform data land in a dashboard automatically. The explanation of what it means almost never is.

Clients rarely need another chart. They need someone to say what changed, why it may have changed, what the team did about it, and what happens next. That’s exactly the kind of writing a language model is well suited to draft — and exactly the kind of writing where an ungoverned model will confidently say the wrong thing.

What is an automated report narrative?

An automated report narrative is a written interpretation of structured performance data generated by software. A reliable system combines deterministic calculations, business context, constrained language generation, and a verification step. It should explain supplied evidence — not infer a convenient story from it.

A production-grade architecture: Python, analytics data, and campaign context

The pattern below is the one I’d recommend for this: a scheduled job pulls current- and prior-period metrics, calculates percentage changes in code, adds completed and planned work from wherever your team tracks it (Jira, Asana, or similar), and sends a compact structured payload to a language model to draft from.

Layer Implementation Purpose
Runtime Python Extract, normalize, and validate metrics
Analytics GA4 Data API + Search Console API Sessions, engagement, key events, clicks, impressions, position
Operational context Project-tracking tool (e.g. Jira) Completed work and the next sprint
Language model A current, supported model API Draft the narrative from the supplied payload
Schedule A scheduled job runner (e.g. GitHub Actions) Run the workflow each reporting period

Model usage cost for a run like this is typically small relative to the analyst time it replaces — comfortably in the range of cents per report, not dollars — though the exact figure depends on payload size and which model you run.

An illustrative shape for the payload:

def build_report_payload(current, prior, context):
    return {
        "metrics": {
            "organic_sessions": metric(current["sessions"], prior["sessions"]),
            "organic_conversions": metric(current["conversions"], prior["conversions"]),
        },
        "completed_work": context["tracked_completed_items"],
        "planned_work": context["tracked_next_period_items"],
        "data_quality_flags": context["tracking_anomalies"],
    }

The important design choice is that the model doesn’t calculate the deltas. Code does. The payload supplies the current value, prior value, percentage change, data-quality flags, and work context as separate fields — which makes the draft easier to test and removes any ambiguity about whether a number rose or fell.

The failure mode this has to guard against

Here’s the risk made concrete with a hypothetical example. Imagine a monthly payload showing:

  • Organic sessions: 12,450 (+4.2%)

  • Organic conversions: 310 (−12.0%)

  • Form submissions: 82 (−22.5%)

An unconstrained model asked to draft an upbeat client update from this could easily lead with the traffic increase and describe performance as “solidifying steady conversion volume” — a sentence that’s fluent, positive, and directly contradicted by the conversion numbers in the same payload. The model wouldn’t have fabricated a number. It would have fabricated the meaning of the numbers — a false interpretation, not a false statistic.

THE OPERATIONAL RISK

A reporting hallucination isn’t limited to a made-up metric. It can also be a false causal claim, an unsupported attribution, or positive framing that contradicts a negative KPI. In a client report, all four are credibility failures.

This is why “ask the model to be accurate” is not a control. Language models produce plausible text from patterns. They don’t independently know which KPI a client considers decisive, whether tracking was stable, or whether a completed task caused a result. Those facts have to be encoded, tested, or withheld.

What stops an AI report from inventing a win?

A longer prompt by itself won’t fix this. The reliable fix is two layers of deterministic control plus a human release gate.

1. Validate direction before generation

A pre-processing step classifies every KPI before the payload reaches the model. When a core metric moves beyond an agreed threshold, the pipeline injects an explicit constraint. Thresholds should be account-specific — a five-percent change may be meaningful for one KPI and noise for another.

def generate_constraints(metric_deltas):
    constraints = []
    for kpi, change in metric_deltas.items():
        if change < -5.0:
            constraints.append(
                f"CRITICAL: {kpi} decreased by {change:.1f}%. "
                "State this decline in the first paragraph. "
                "Do not describe it as positive, stable or optimal."
            )
    return "\n".join(constraints)

2. Constrain claims and attribution

The system instruction should require exact values, separate completed work from proven causes, and define a failure response for missing or anomalous data. Completed work can show that something happened; it can’t by itself prove that it caused a result.

REPORTING RULES
  1. Every quantitative claim must quote a value supplied in the JSON payload.

  2. State material negative changes plainly; do not soften or reverse their direction.

  3. Do not claim causation unless the payload contains an approved evidence note.

  4. Put uncorrelated completed tasks under “Work completed,” not “Drivers of performance.”

  5. If required data is missing or anomalous, return: [DATA ANOMALY: MANUAL REVIEW].

3. Reconcile the draft after generation

A post-generation validator should extract percentages and metric names from the draft, compare them with the source payload, and scan for prohibited positive framing near negative deltas. It should also check that each material decline appears in the narrative. A mismatch should fail the run rather than trigger an automatic rewrite loop that could quietly hide the error.

Control Pass condition Failure action
Numeric reconciliation Every stated number exists in the payload Stop and review
Directional consistency Language matches positive/negative direction Stop and review
Material KPI coverage Every threshold breach is disclosed Stop and review
Attribution discipline Causal claims have approved evidence Remove claim or review
Data completeness Required fields and tracking flags present Use anomaly response

4. Keep a human release gate

The final human review shouldn’t be ceremonial. Someone checks the headline conclusion, data-quality flags, causal language, and next actions before publication. The automation owns assembly and first-draft prose. The account team keeps accountability for interpretation and client communication.

What this approach buys you

Done well, a pipeline like this turns report narrative drafting from one of the more repetitive parts of the reporting cycle into a short review step. The actual time saved depends on account volume and how much of the current process is already structured — but the larger gain isn’t really about speed. It’s role clarity: staff stop spending the reporting cycle assembling sentences and spend it on exceptions, interpretation, and strategy instead.

A reusable framework for AI-generated performance commentary

  1. Define the decision metrics. Identify the KPIs that change the client conversation and set materiality thresholds for each one.

  2. Calculate outside the model. Normalize date ranges, percentages, and comparison logic in code — not in the prompt.

  3. Separate facts from context. Store metrics, completed work, planned work, tracking anomalies, and approved causal notes as distinct fields.

  4. Write a claim policy. Require exact values, plain treatment of declines, and explicit handling of missing data.

  5. Validate the output. Reconcile numbers, direction, KPI coverage, and causal language against the payload.

  6. Escalate instead of improvising. A failed validation should create a review task, not invite the model to keep guessing.

  7. Log the evidence. Retain the input, model/version, prompt version, validator result, reviewer, and publication timestamp.

RULE OF THUMB

Use an LLM for language transformation, not for arithmetic, source-of-truth selection, or final accountability. The more consequential the report, the more of the pipeline should remain deterministic.

Model versions are dependencies

Anyone building this needs to treat the model itself as a versioned production dependency, not a fixed constant. As one concrete example of how fast this moves: Anthropic deprecated claude-3-5-sonnet-20241022 on August 13, 2025, and retired it on October 28, 2025. A pipeline built against a specific model ID needs a plan for what happens when that model is retired — pin the model ID explicitly, and rerun your evaluation set before deploying a replacement. A model swap can change tone, instruction-following, extraction behaviour, and cost even when the prompt is unchanged.

Keep a small test suite that includes a normal month, a strong decline, mixed metrics, missing data, a tracking anomaly, and an unsupported causal claim — and rerun it every time the underlying model changes.

Frequently asked questions

Can AI write a monthly marketing report narrative?

Yes. An LLM can turn structured performance data and campaign context into a readable first draft. For client-facing use, calculations should happen in code, claims should be constrained to supplied evidence, and the draft should pass automated validation plus human approval.

How do you prevent hallucinations in automated reports?

Don’t rely on prompting alone. Reconcile every number with the source payload, test whether the language matches the metric direction, require material declines to be disclosed, prohibit unsupported causal attribution, and stop the workflow when data is incomplete or anomalous.

Why not use a dashboard tool’s built-in narrative features?

A widget-level summary may not have access to historical account context, completed work, deployment notes, or custom data-quality flags. A custom pipeline can unify those sources and apply account-specific thresholds before language generation.

Can the model recommend what to do next month?

Yes, when planned tasks and decision rules are included as structured context. Recommendations should be labelled as planned actions — not presented as proven consequences of last month’s performance.

What happens when the GA4 Data API reaches a quota?

Use bounded retries with backoff, monitor property quota consumption, and fall back only to a clearly timestamped cache. Google’s Data API can return property quota details when requested. A stale fallback should be disclosed rather than silently treated as current data.

Does this eliminate the need for account managers?

No. It reduces repetitive drafting. Account managers still validate interpretation, explain uncertainty, discuss priorities, and own the client relationship.

Should FAQ schema be added to this article?

Not for the purpose of earning a Google FAQ rich result — Google deprecated that feature in 2026. Keep the FAQ because it answers real questions clearly; use Article and Person/Organization markup to describe the page and its author accurately.

The lesson: automate the draft, not the responsibility

Automating executive commentary isn’t about removing human accountability. It’s about moving the human from blank-page writer to reviewer, exception handler, and strategist.

A system like this can optimize for speed and fluent prose, or it can optimize for traceability — every important sentence pointing back to a metric, an approved context field, or a clearly labelled plan. Only the second version is safe to hand a client. That’s what makes the time savings usable without spending client trust.

For the broader operating philosophy behind this approach, read “There Are No 7 Prompts: 50 Things I Actually Do With AI Inside a Working Agency”. For the reporting foundation, see “The Client Dashboard That Replaced a Monthly Reporting Call”.

Sources and methodology

Product-status and platform claims in this article were checked against the following first-party documentation:

  • Google Analytics Data API overview

  • Google Analytics Data API limits and quotas

  • Anthropic model deprecations

  • Google: creating helpful, reliable, people-first content

  • Google: generative AI content guidance

  • Google Search documentation updates


About David Henderson

David Henderson is an SEO and digital marketing practitioner with more than 25 years of experience building, managing, and evaluating search, reporting, and content systems. His current work focuses on practical AI integration: where automation improves production, where it fails, and what still requires human judgment. He is the founder of Unwired Web Solutions and writes at davidhenderson.ca.