The short answer
- Yes, the direction is real. Rigorous field evidence shows AI assistance makes support staff measurably more productive[3], and analysts forecast most routine service issues will be resolved autonomously by 2029[1].
- No, the shortcut isn’t. Reliability, quality and accountability are where programmes fail[2],[4],[10],[11]. The winners treat agents as a governed system, not a chatbot swap.
- So it is worth it if you start with assistance, automate narrow and safe intents, and build context, controls and measurement first. It is not worth it as a pure headcount play.
1 · The journey
From scripts to agents that act
Customer service automation has moved through four recognisable stages. Most organisations are somewhere between the second and third. The difference between them is not the model. It is whether the system can take an action and be held to account for it.
Anthropic draws a useful line: workflows orchestrate models and tools through predefined code paths, while agents let the model direct its own process and tool use[7]. The same guidance is to find the simplest solution that works and add complexity only when needed. That advice is the quiet backbone of this article.
Scripted
Menus, decision trees and FAQ bots. Predictable, cheap, and brittle the moment a customer goes off script.
Assisted
A model drafts replies and surfaces knowledge while a person decides. This is where the strongest peer-reviewed evidence sits today.
Agentic
Agents plan, call tools and complete the task: refund, rebook, update the record. Humans handle exceptions and approvals.
Networked
Your agents deal with your customers’ agents. Identity, authority and audit become the product, not an afterthought.
2 · The architecture
What an agentified CX stack actually contains
An agent is a loop: reason, act with a tool, observe the result, repeat. The ReAct paper made that pattern explicit[5], and it underlies most tool-using agents today. Making that loop coherent over time takes more than a prompt. Research on generative agents showed memory, reflection and planning working together as the architecture that keeps behaviour consistent[6].
In customer experience that translates into six layers. Skip one and the failure shows up in front of a customer.
Channels
Chat, voice, email, messaging and in-product. One conversation, many surfaces.
Conductor
Routes work to the right persona or specialist agent and keeps the plan, instead of one model improvising everything.
Context and memory
The customer’s history, entitlements, policies and prior decisions, permissioned, so an agent answers from facts not guesses.
Tools
Governed access to billing, orders, CRM and knowledge through open protocols such as MCP.
Policy and approval
Limits an agent cannot exceed, and human sign-off for anything irreversible, costly or sensitive.
Evaluation and observability
Traces, reliability testing and outcome measures, so you find failures before customers do.
Two open standards are making the tool and collaboration layers less proprietary. Anthropic donated the Model Context Protocol to the Linux Foundation’s Agentic AI Foundation in December 2025[8], and Google’s Agent2Agent protocol moved to the Linux Foundation as well[9]. MCP connects an agent to tools and data; A2A lets independent agents talk to each other. Betting on open protocols is how you avoid rebuilding the integration layer every time the model market shifts.
3 · Is it worth it?
The evidence, both ways
The honest answer lives in the tension between two kinds of evidence: strong results where AI helps people, and sobering results where AI is left to act alone.
What supports the case
≈15%
more issues resolved per hour[3]
In a study of 5,172 support agents, an AI assistant lifted productivity by about 15% on average, with the largest gains for less experienced staff (the 2023 working paper reported 14% overall and 34% for novices).
80% / 30%
Gartner’s 2029 forecast[1]
Gartner predicts agentic AI will autonomously resolve 80% of common customer service issues by 2029, with a 30% reduction in operational costs. This is a forecast, not a measurement.
What should give you pause
<50%
task success in a service benchmark[4]
On τ-bench, strong function-calling agents completed fewer than half of simulated retail and airline tasks, and in retail succeeded on all eight repeated attempts of the same task less than a quarter of the time.
>40%
of agentic projects cancelled by 2027[2]
Gartner expects cancellations over escalating cost, unclear value and weak risk controls, and warns of “agent washing”: old chatbots and RPA rebranded as agents.
Liability
you own what your bot says[10]
In Moffatt v Air Canada a tribunal rejected the airline’s argument that its chatbot was a separate entity. A small-claims ruling, but a clear signal.
Quality
cost was over-weighted[11]
Klarna, an early and vocal adopter, said it moved back towards human support because optimising for cost had produced lower quality. This is the CEO’s account as reported.
A widely quoted MIT NANDA report found most generative AI pilots delivered no measurable profit-and-loss impact[12]. Its definition of success and sample have been criticised, so we treat it as direction rather than a rate. It does point the same way as the rest of the evidence: value comes from integration into real workflows, not from a demo.
Our reading. The gains are real where AI amplifies people inside a governed process, and fragile where it is asked to act alone without context, limits or measurement. The cancellations Gartner expects[2] are a forecast about governance and value discipline, not about the technology failing to work.
Try your own numbers
A break-even sketch with a deliberate quality haircut. Move the sliders to see what share of contacts agents must resolve before this pays for itself.
Monthly result
+$34,200
$43,200 of handling cost avoided, less $9,000 to run it.
Break-even: about 6% of contacts resolved by agents at these costs.
On your numbers the programme pays for itself. The next question is whether you can hold quality while you get there.
Illustrative arithmetic, not a benchmark. Defaults are placeholders, not Neocortex customer results. It ignores revenue effects, one-off build cost and the value of faster resolution, all of which matter in a real business case.
4 · Where this is heading
How the journey will evolve
Forecasts are not facts, and the dates below are Gartner’s horizons[1],[2] plus our own judgement. The shape matters more than the year: assistance first, bounded authority next, then agents serving other agents.
Now to 2027
Assist first, automate the narrow and safe
- Agent assist for every human conversation; measure handle time and resolution quality.
- Fully automate only high-volume, low-risk, well-policied intents (status, simple changes).
- Stand up the context layer and the evaluation harness before you scale volume.
2027 to 2028
Agents with authority, bounded
- Agents take actions with spending and scope limits; humans approve the irreversible.
- Reliability measured as “every time”, not “usually”, because customers experience the worst case.
- Cancelled projects are the ones without value tracking or risk controls. Don’t be one.
2028 to 2029+
Agent-to-agent service
- Customers’ own agents negotiate with yours; Gartner calls these “machine customers”.
- Identity, delegated authority and audit trails become the competitive layer.
- Open protocols (MCP for tools, A2A between agents) reduce lock-in as the estate gets more complex.
5 · How Neocortex gets you there
The governed route, one layer at a time
Neocortex is built around the same layers the evidence says you need, with the governance in the execution path rather than bolted on afterwards.
- Conductor and personas. A conductor persona plans and delegates to specialist personas and a reviewer, so one model is not improvising everything. See the avatar builder.
- The Enterprise AI Brain. Customer, policy and decision context lives in a governed layer with permissions, so agents answer from facts. Read the white paper.
- Tools through open protocols. Connect the systems you already run instead of replacing them; see integrations.
- Approval and audit. Humans sign off the irreversible and the sensitive, and every action leaves a record of who or what did it. See security.
- Australian hosting, or your own cloud. Data residency is a design decision from day one.
- Start small, compound. A focused discovery, a few priority intents, weekly delivery, measured against the break-even above.
We make no customer-outcome claims on this page. The figures above are other people’s research, cited so you can check them. The value of a pilot is that it produces your own.
Scope your first agentic CX pilot
Get a proposal, or talk it through with the team.
References
Academic papers and primary analyst releases are listed first-hand where we could read them. Where we rely on press coverage of a primary source, the entry says so.
[1] Analyst
Gartner (5 March 2025). Gartner Predicts Agentic AI Will Autonomously Resolve 80% of Common Customer Service Issues Without Human Intervention by 2029.
Forecast: 80% of common service issues resolved autonomously and a 30% reduction in operational costs by 2029.
https://www.gartner.com/en/newsroom/press-releases/2025-03-05-gartner-predicts-agentic-ai-will-autonomously-resolve-80-percent-of-common-customer-service-issues-without-human-intervention-by-20290[2] Analyst
Gartner (25 June 2025). Gartner Predicts Over 40% of Agentic AI Projects Will Be Canceled by End of 2027.
Forecast: over 40% of agentic projects cancelled over cost, unclear value or weak risk controls; "agent washing" by vendors.
https://www.gartner.com/en/newsroom/press-releases/2025-06-25-gartner-predicts-over-40-percent-of-agentic-ai-projects-will-be-canceled-by-end-of-2027[3] Academic
Brynjolfsson, E., Li, D. & Raymond, L. (2025). Generative AI at Work. The Quarterly Journal of Economics, 140(2), 889–942. doi:10.1093/qje/qjae044 (preprint arXiv:2304.11771).
Field study of 5,172 support agents: about 15% more issues resolved per hour, with the largest gains for less experienced workers.
https://arxiv.org/abs/2304.11771[4] Academic
Yao, S., Shinn, N., Razavi, P. & Narasimhan, K. (2025). τ-bench: A Benchmark for Tool-Agent-User Interaction in Real-World Domains. ICLR 2025 (arXiv:2406.12045).
Even strong function-calling agents completed under half of simulated retail and airline service tasks, and were inconsistent across repeated trials.
https://arxiv.org/abs/2406.12045[5] Academic
Yao, S. et al. (2023). ReAct: Synergizing Reasoning and Acting in Language Models. ICLR 2023 (arXiv:2210.03629).
The reason-act loop that underpins most tool-using agents.
https://arxiv.org/abs/2210.03629[6] Academic
Park, J. S. et al. (2023). Generative Agents: Interactive Simulacra of Human Behavior. UIST 2023 (arXiv:2304.03442).
Memory, reflection and planning as the architecture that makes agent behaviour coherent over time.
https://arxiv.org/abs/2304.03442[7] Industry
Anthropic (19 December 2024). Building effective agents.
Workflows versus agents, and the advice to use the simplest solution that works.
https://www.anthropic.com/engineering/building-effective-agents[8] Industry
Anthropic (9 December 2025). Donating the Model Context Protocol and establishing the Agentic AI Foundation.
MCP moved to vendor-neutral governance under the Linux Foundation.
https://www.anthropic.com/news/donating-the-model-context-protocol-and-establishing-of-the-agentic-ai-foundation[9] News
SD Times (2025). Google’s Agent2Agent protocol finds new home at the Linux Foundation.
A2A, the agent-to-agent protocol, moves to open governance.
https://sdtimes.com/ai/googles-agent2agent-protocol-finds-new-home-at-the-linux-foundation/[10] Legal
Moffatt v. Air Canada, 2024 BCCRT 149 (Civil Resolution Tribunal of British Columbia, 14 February 2024).
A company was held responsible for what its website chatbot told a customer. A small-claims decision, not binding precedent, but a clear signal.
https://www.canlii.org/en/bc/bccrt/doc/2024/2024bccrt149/2024bccrt149.html[11] News
Entrepreneur (2025). Klarna Is Hiring Customer Service Agents After AI Couldn’t Cut It on Calls, According to the Company’s CEO (reporting Bloomberg’s interview with Sebastian Siemiatkowski).
A high-profile automation leader rebalanced towards human support after concluding cost had been over-weighted and quality had suffered.
https://www.entrepreneur.com/business-news/klarna-ceo-reverses-course-by-hiring-more-humans-not-ai/491396[12] News
Fortune (18 August 2025). MIT report: 95% of generative AI pilots at companies are failing (reporting MIT NANDA, The GenAI Divide: State of AI in Business 2025).
Most pilots showed no measurable P&L impact in the study window. The definition of success and sample have been criticised, so read it as direction, not a precise rate.
https://fortune.com/2025/08/18/mit-report-95-percent-generative-ai-pilots-at-companies-failing-cfo
Figures were checked against the sources in October 2026. Gartner’s releases are forecasts. The productivity result is from one large company’s customer support operation and may not transfer to every setting.
