NEW  Recursive self-improvement is live, your agents build, test and improve themselves in production, every week. See how →
Recursive self-improvement · RSI

Recursive auto-AGENTS™When agents build their own agents.

Recursive self-improvement (RSI) is the point where an AI system helps design, test and improve the next one. In the contact center, it means agents that build other agents from your own material, measure themselves against real outcomes, and get better every week, with a human setting direction and holding the guardrails.

0
Prompts to hand-write
0
Conversations become signal
auto
Evals generated & run
0
Improvement cadence
Recursive self-improvement

For most of the contact center's history, humans built every step. Someone wrote each intent, scripted each reply, and hand-tuned every prompt. Improvement was a project, scoped, staffed, and shipped once, then left to age.

Recursive auto-AGENTS™ changes the unit of work. Agents now build other agents from your existing material, test them against evals they wrote themselves, and improve them from real outcomes, while humans set direction and hold the guardrails. Improvement stops being a project and becomes a property of the system.

This page explains how the loop works, why it was hard to make safe, and where a human stays in it, on purpose.

How we got here

From hand-built flows to a loop that closes.

2016–2019

Hand-built flows

Every intent and reply scripted by hand. A change meant a ticket and a release.

2020–2023

Assisted authoring

Teams drafted snippets with LLMs, then copied, pasted and tuned them by hand.

2024–2025

Agents that build

Agents composed whole flows from documents; humans reviewed everything before launch.

Today

The loop closes

Agents build, test and improve other agents in production. Humans set direction and approve.

Next

Wider autonomy

More of the loop runs unattended as evals and guardrails earn trust, never all of it.

The loop

Resolve, measure, learn, optimize, on repeat.

Recursive self-improvement (RSI) runs a closed loop in production. Every interaction is resolved, measured by Auto QA, learned from, and turned into a concrete change to prompts, routing, tools and guardrails, then the loop starts again. Every change is gated and reversible.

01

Resolve

AI and human agents handle live conversations end to end across voice and chat.

02

Measure

Auto QA grades 100% of interactions against your rubric, quality, compliance, outcome.

03

Learn

Winning patterns are extracted from top AI and human conversations; gaps are diagnosed.

04

Optimize

Prompts, routing and guardrails retune in production, gated, measured, reversible.

agent on a live customer call
Live · voice & chatEvery conversation feeds the loop
support specialist
Human agentsBest conversations, learned
operations team reviewing performance
In productionImproving week over week
What the loop does

Five jobs, one continuous loop.

From standing up a brand-new agent to rescuing a live call, Recursive auto-AGENTS™ does the work that normally needs a prompt engineer, a QA team and an eval harness, continuously, and with a human in the loop.

No prompt engineering required.

01 · Build agents from your SOPs and docs

Point it at your standard operating procedures, help-center articles, policies and past transcripts. It composes a working auto-AGENT, flows, tools, knowledge and guardrails, that you refine in plain language, not prompt syntax. What took weeks of prompt building becomes a first draft in minutes.

SOPs
docs · transcripts · KB
0
prompts written
Compose agent · from sourcesdraft
SOPBilling & refunds SOP.pdfread
KBHelp center · 214 articlesindexed
TR12,400 past transcriptslearned
Drafting flows, tools & guardrailsbuilding
Billing agent, ready to reviewdraft
no-prompt Refine in plain language, describe the change, not the prompt.
Learns from AI and human agents alike.

02 · Improve from Auto QA and past conversations

Every Auto QA score and every past conversation, from your AI agents and your best human agents, becomes training signal. Recursive auto-AGENTS™ promotes the patterns that resolve faster and score higher, so agents improve week over week without a manual tuning cycle.

100%
conversations used
AI + human
both sources
Learning signallast 7 days
Resolution rate↑ 6
QA score↑ 4
Patterns promoted31
promoted Winning phrasing from top human agents, applied to the AI.
It tests your agents before your customers do.

03 · Auto-test and generate evals

It writes eval suites from your requirements and edge cases, then places live test calls to agents still under development, probing tricky paths, grading each response and flagging regressions. Agents ship measured, not hoped-for, and every release is gated against the evals.

auto
evals generated
24/7
test calls
Eval run · agent v14 (dev)nightly
Refund past 30 days, edge casepass
Angry caller · de-escalationpass
Partial payment + promo stackfail
Spanish ↔ English handoffpass
gated 142/145 evals pass · 3 regressions blocked from release.
The same engine lifts your people.

04 · Improve human conversations

Recursive auto-AGENTS™ reads human interactions the same way it reads AI ones, surfacing what top performers do, and turning it into personalized coaching agendas and in-the-moment guidance. The whole floor moves toward the best, not the average.

Coaching · auto-generatedagent · A. Rivera
1Acknowledge frustration earlier, great recovery at 02:14empathy
2Offer the loyalty tier when renewal comes upupsell
3Swap internal jargon for plain languageclarity
from the best Modeled on top-scoring conversations across the team.
Improves the outcome before the call ends.

05 · Intervene mid-conversation

When a live conversation is heading the wrong way, Recursive auto-AGENTS™ steps in in real time, offering the next best action, a correction or a better path to whoever is on the call, human or AI. It doesn't just learn for next time; it improves this outcome, now.

<500ms
real-time nudge
live
human or AI
Live call · billing02:41 · in progress
Caller"This is the third time I've called about the same charge."
Agent"I understand, let me pull up the history."
Recursive auto-AGENTS™ · intervene
Repeat contact detected. Lead with an apology + same-day credit; skip re-verification, identity confirmed on call #2.
Why this was hard

Real self-improvement, with a human still in the loop.

The simplest way to make a system improve itself is to let it change itself freely. In a contact center you can't. Every conversation is regulated, brand-sensitive and often someone's worst day. An agent that silently rewrites its own behavior is a compliance and trust problem, not a feature.

But keeping a human in the loop the naïve way, approving every change by hand, throttles improvement back to the speed of a review queue. That is the core tension: oversight wants to slow things down; recursion wants to speed them up.

We resolved it by changing what the human does, not removing them. The loop does the labor, proposing changes, writing and running evals, measuring live outcomes. The human does the judgment, setting the goal, defining the guardrails, and approving changes that clear the evals. Every change is gated, measured and reversible, so autonomy grows exactly as fast as the evidence earns trust.

Think of it like a bike: the loop coasts on its own, but a human has to pedal in between. That pedal, a quick human approval between the automated runs, is what keeps recursive self-improvement safe.
food_ordering_bot · loopv4.0
Run & testauto-AGENT runs 12 caller scenarios
auto
Human approvesreview the fixes, pedal to continue
you
Rebuild & shipfix applied · gated & reversible
auto
The loop coasts on its own, a human has to pedal in the middle
Your IP · your environment

The improved agents are your company's IP, and they stay on your premises.

Every agent Recursive auto-AGENTS™ builds, and everything it learns from your conversations, belongs to you and stays inside your environment, run it in your own cloud, private VPC, or fully on-prem. Your data and the improvements it produces are yours alone: never shared, and never used to train anyone else's models.

On-prem / VPCYour data stays inNot used to train othersSOC 2 Type II · HIPAA · PCI · GDPR
Evidence

What happens when the loop runs.

Two views from production: how quickly changes start shipping without human edits as the evals earn trust, and a before/after from a 30-day pilot.

Share of agent updates shipped without human edits
Weekly · single deployment · higher = more autonomy earned
0 20 40 60 W1W2W3W4W5W6W7W8 8% 64%
How to read this: a change ships "without human edits" when it clears the auto-generated evals and the reviewer approves it unchanged. Early on, almost everything is edited; as the evals prove reliable, more clears untouched, autonomy the system earns, not autonomy assumed.
30-day production pilot · anonymized · 420-agent contact center
MetricBeforeWith Recursive auto-AGENTS™
Autonomous resolution rate61%78%
Auto QA score (0–100)7486
Prompts hand-written per agent~400
Time to stand up a new agent3–5 weeks2 days
Change → productiondays, manualcontinuous, gated
Human edits per 100 changes10012
Directional results from a single deployment; your numbers depend on baseline quality, volume and how much autonomy you enable. Every change remained gated and reversible throughout.
Beyond an agent-builder

Most tools stop at “build.” This one never stops.

An agent-builder helps you ship an AI agent, then hands it off. Recursive auto-AGENTS™ does everything they do, build from your docs, auto-test edge cases, learn from your best transcripts, then keeps going: it lifts your human agents, steps into live calls, and self-tunes in production, week after week.

Capability
Typical agent-builder
Recursive auto-AGENTS™
Build agents from SOPs & docs, no prompt engineering
Auto-generate evals, test edge cases, fix on failure
Learn from your best transcripts & recordings
Improves your human agents, too
Intervenes live, mid-conversation
Self-tunes in production, continuously, not just at build time
Voice-native, sub-500ms, not just chat
A tenth of the inference cost, via TokenTrim™
Why it matters

Improvement stops being a project. It becomes a property.

Days

Launch, not build

New agents drafted from your existing SOPs and docs in minutes, refined in plain language, live in days instead of weeks of prompt engineering.

0

Quality that compounds

Every conversation and Auto QA score feeds the loop, so resolution and quality climb week over week, no manual tuning cycle.

Safe

Gated and reversible

Auto-generated evals gate every release, and each production change is measured and one click from rollback. Momentum without risk.

"We stopped hand-writing prompts and hand-picking calls to review. It builds the agent from our own playbooks, tests it overnight, and quietly gets better every week, including our human team."
VP, Customer Operations
Enterprise retailer · 900 agents
Proof

From your playbook to a self-improving agent.

Bring your SOPs and a month of conversations. We'll draft an agent, generate its evals, and show the loop lifting both AI and human performance, on your systems, in your languages.

See it live

Hand us your playbook. We'll hand back an agent that improves itself.

In 30 minutes we'll draft an agent from your docs, generate its evals, and show the recursive loop at work, for your AI and your people.