- The problem: every client question travelled the chain Client → PM → Specialist → PM → Client. Answers took days, and the team was looking at channels rather than at the client's business. The result: falling NRR and real churn.
- The solution: The Forge, an internal AI platform built on Claude Code. Not a "chat with AI", but a library of skills that solve specific problems, fed with the full, daily-refreshed client context: GA4, Search Console, Google Ads, Meta Ads, the project CRM and call transcripts.
- The results: monthly reports down from 2 hours to about 15 minutes. Keyword research from 2 days to an hour. The PM evolving from middleman into strategist. NRR recovering from a low of about 59% to about 80%, alongside a rise in structured reports from roughly 20 to 75 per month.
A note on the screenshots: The Forge's interface is in Polish, since that is the language the team works in. They are otherwise unedited, except that email addresses have been covered over and marked e-mail ukryty ("email hidden").
The problem: an agency that looks at the channel, not the business
A typical situation in a marketing agency. The client asks: "Why did sales drop last month?" The question reaches the PM. The PM does not hold complete knowledge of Google Ads, SEO and analytics at the same time, because nobody does. So they go looking for a specialist. The specialist has a calendar booked two days out. The answer comes back to the PM, and the PM explains it to the client.
The chain Client → PM → Specialist → PM → Client closed in a day or two. Our PMs took the communication on the chin and made sure the client stayed informed, but these days even two days is too slow. On top of that, the PM had no room to ask questions when they actually needed to. They always waited for the specialist to find a slot.
And the problem ran deeper than speed:
"What really irritated me was how often specialists and PMs look at the channel rather than at the client's business. I decided we needed a space that would demand business quality, not channel quality."
Everyone did their own patch well. Campaigns were optimised, rankings were climbing. But nobody was assembling it into an answer to the question the client is actually asking: what does this mean for my business, and what do we do next? The team was productive, but not always effective: doing a lot, not necessarily the things that mattered most.
The value the agency genuinely delivered was not being presented properly. And value a client cannot see does not exist for that client. We could see it in hard numbers: NRR for the mature project cohort slid from about 78% to a low of about 59%.
The decision: build, but not a "platform for everything"
Before The Forge existed, we tested an approach on ourselves that did not work. In the first iteration of a parallel project, people were handed an open tool in which they could build whatever they liked. That turned out to be the barrier: the entry threshold was too high, the procedures too complicated, and the resistance entirely natural.
"Adopting AI in an organisation has to start by levelling the entry threshold on the human side. Otherwise it will not land: social resistance will kill the project, or breaking through it will take a very long time."
So The Forge was built on the opposite premise: people do not adopt platforms, they adopt solutions to their problems. Instead of an open AI chat, the team got a library of ready-made skills: generate the monthly report, run keyword research, prepare a Google Ads negative list, build a strategic dashboard. Each skill solved a specific, familiar pain. And along the way it got people comfortable working with aggregated data, building the trust needed for the next step.
We chose Claude Code as the foundation. Three things decided it: managing the whole tool through a single central repository, so everyone in the company works on the same current version; the strongest models in benchmarks; and native support for orchestrated skills. The model simply understands multi-step procedures well, so systems built around it stay easy to maintain. The alternatives, even those competitive in benchmarks, did not offer the same management toolkit.
How it works: the anatomy of a single answer
Rather than stay at the level of generalities, let us walk through what happens when a PM asks The Forge the same question that used to trigger a two-day relay: "Why did sales drop last month?"
1. Ten-second sign-in. The PM opens The Forge and enters an email address and a one-time code from an authenticator app. A code rather than a password, deliberately: the code expires in moments, so even if it ends up saved in a chat history it is worthless. We treat client data security here without compromise.
2. The context loads itself. At the start of a session the platform pulls the full project context from our API: tasks and progress from the project CRM, the PM's strategic notes, transcripts of recent client calls, SEO and analytics data. Automations refresh that data daily, so the context is always current. The PM does not have to paste anything in or explain to the tool "what this project is about".
3. The question is routed to the right sources. A decision map lives in the platform repository: question type → data source. A question about a sales drop is a compound one, so The Forge reaches for a combination of sources at once:
- per-product e-commerce data — what exactly fell: which product, which category, at which stage of the funnel;
- traffic sources and channels — whether the drop came from organic, paid or direct;
- daily Google Ads campaign data — whether the ads delivered: cost, conversions, impression share, visibility lost to budget or rank;
- call transcripts and the PM's notes — whether a decision in the background explains it: an offer change, a paused budget, seasonality the client mentioned in a meeting.
4. Headline numbers come from the source of truth. Whole-store session and revenue totals are always taken from a dedicated endpoint holding monthly totals, never by summing dimension tables, since those inherently fail to reconcile in analytics APIs. That sounds like a detail, but details like it decide whether the numbers in the answer match the GA4 panel the client is looking at.
5. Analysis runs on raw rows; conclusions sit with the model. The API deliberately returns no ready-made conclusions, only raw daily or monthly rows shaped for the question at hand — plus, wherever possible, deterministically built context: data processed by algorithms rather than by the model's interpretation. The model aggregates, compares periods, calculates shares and stitches that together with qualitative context from calls. If data for a period is missing, it says so plainly. Fabrication is banned at the architectural level: no records means a "no data" message, not an invitation to hallucinate.
6. The answer is about business, not channels. The PM gets not a table but an answer along the lines of: sales fell by X% mainly in category Y; paid traffic delivered, but organic lost visibility on product queries; and the transcript of the last meeting shows the client cut the budget for Z. Plus what is worth doing next. Which is exactly what the PM used to spend two days chasing across three people.
The whole path takes minutes. And, just as importantly, the PM can keep digging. The specialist's calendar stopped being the bottleneck on curiosity.
Skills: where the answer should be a standard, not an improvisation
Free-form conversation handles ad hoc questions. But part of agency work is repeatable processes where what counts is not only accuracy but unchanging, auditable quality. That is what skills are for: packaged, multi-step procedures launched with a single command.
Our flagship reporting skill shows it best. Under the hood it looks like this:
A deterministic engine for the numbers. A script pulls data from seven sources in parallel (a one-year window for trends, with automatic retries on errors), and then a separate, deterministic engine calculates everything calculable: reconciling totals, splitting brand from acquisition traffic, quantifying changes, setting priorities. One rule is iron: no number in the report originates in a language model. The model receives finished, calculated values and is responsible only for the narrative and the business framing. That is why the report always agrees with the panels the client has access to.
Tone Card: a report written in the PM's voice. Before the report is generated, a separate agent analyses that PM's best-rated historical reports and distils a style profile from them: how this person opens and closes, which phrases they use, how they balance bad news, whether they state recommendations outright or propose them. The client receives a report that sounds like their PM, not like a content generator. The profile carries style only; carrying over content or numbers from old reports is forbidden.
Continuity Brief: memory of the conversation with the client. A second agent extracts commitments and threads from the previous report: what we promised to check, what we have already explained, what story we have been telling for months. The new report continues that narrative, accounts for the promises, and does not sell the same curiosity as a fresh discovery for the third time. The monthly reset — the besetting sin of most agency reporting — disappears.
An expert in the loop before anything reaches production. No skill goes into use without sign-off from a domain expert. The keyword research skill was built in a loop with our lead SEO specialist: he assessed output quality, and without his approval the skill did not ship. When it did, the first user summed it up like this:
"Two days of work turned into an hour."
The same philosophy governs the analytical skills: the QBR skill (quarterly business review) has three validation gates built in — facts, quality, design — and the rule that "maths happens only in scripts", with data-quality flags and gaps surfaced explicitly in the document footer. Transparency instead of sugar-coating.
From user to author
The most interesting part began a few weeks after launch. When somebody works out something valuable in a free-form conversation with The Forge, all they have to write is: "turn this into a skill". A one-off action becomes a reusable standard: available on every project, versioned, shareable and rateable. There is also a full skill builder that walks you through the process step by step, from the name and data sources to publication in the company library.


Today the catalogue holds, alongside the "factory" skills, skills built by specialists on the team: a Google Ads placement report from our PPC specialist, proposal builders from the sales team, and more. Each tagged with its author, version and rating. The loop closed: a tool meant to get people comfortable with AI is now being extended by them.
Results: numbers and a change of roles
Scale about four months after the decision to build:
- 43 skills in the catalogue: 13 built into the platform and 30 built by the team, including skills written by users themselves
- 28 of them had recorded usage in the May–August 2026 window
- 25 monthly active users — 33.8% adoption, that is 25 of 74 accounts (as of the end of August 2026)
- monthly report: from 2 hours to about 15 minutes
- keyword research: from 2 days to about 1 hour

But the most important chart hangs elsewhere: NRR for the mature cohort set against the number of reports generated.

NRR had been sliding from about 78% to a low around 59–60%. Exactly in the period when the number of structured reports rose from roughly 20 to 75 per month, the trend reversed and NRR returned to about 80%. To be straight about it: this is a correlation on anecdotal evidence, not a controlled experiment — we are comparing a cohort against point-in-time periods. We did check that there were no anomalies at the start of the cohort that would explain the recovery. And the mechanism behind it is hard to ignore: clients started staying once they began regularly seeing the value they were getting, and the plan for where we were heading.
The change that does not show up in the numbers: the PM stopped being a forwarding manager. They can ask questions whenever they want, without waiting for a slot in a specialist's calendar, and they are evolving into a marketing strategist able to challenge a specialist on the merits on the way to the best solution for the client. The specialists were not replaced: they were relieved of reproductive work and pulled into the quality loop as reviewers of skills in their own fields. Both the speed and the quality of the old relay improved beyond measure.
Adoption as a KPI, not a pious hope
We are still pushing the tool out across the organisation, and we treat adoption as a product metric: measured monthly as the share of accounts with at least one skill use, against an explicit target of 85% of the organisation within 6 months, with milestones along the way.

The roadmap for the tool itself is equally concrete: The Forge should be not merely a necessary asset but above all one people demand. A partner that helps deliver value to clients, and helps specialists and PMs grow and work comfortably, removing blockers that cannot be cleared without it or that cost far too much time.
What this rollout taught us
- Start from problems, not from a platform. People adopt solutions to their own pain. The creative mode arrives on its own, once there is trust.
- Demand business quality, not channel quality. The tool has to assemble channels into an answer to the client's question about their business. Otherwise you are building a faster version of the old problem.
- Feed the model context, not raw dumps. A daily-refreshed aggregation of data from every system, plus transcripts of real client conversations, is the difference between an answer and a guess.
- Calculate numbers deterministically, outside the model. The language model is responsible for narrative and framing; every number comes from a script or from the data. That is the precondition for trust — the team's and the client's.
- An expert in the loop. No skill reaches production without a domain expert's sign-off. That is what separates a standard from a content generator.
- Measure adoption and set targets for it. Adoption is a product KPI, with a goal, milestones and a dashboard. Not "let us see how it catches on".