Developer API · Headless & API-first

Build on the engine.

Every capability is an API. Run auto-AGENTS™ from your own stack with streaming inference, MCP-native tool use and 170+ connectors, no required UI. Your agents call ours; ours call your tools. Model-agnostic, low-latency, and deployable in cloud, VPC or on-prem.

Powered by streaming inference sub-500ms & MCP-native tools 170+ connectors

0
Native MCP connectors
0
Streaming first-token
0
Concurrent sessions
0
Compliance regimes
Headless · API-first

Every capability is an endpoint.

There is no UI you must adopt. Agents, runs, tools, knowledge, voice, evals and analytics are all exposed as a clean REST + streaming API. Drive auto-AGENTS™ from your own application, your IVR, or your own agents, the engine is headless by design.

/agentsCreate & run agentssync or token-streamed, voice or text
/toolsRegister MCP toolsyour tools exposed to every agent
/knowledgeIndex & retrieveupload, embed, test retrieval over API
/runsInspect every runtraces, tool calls, scores, replay
MCP-native tool use

Your agents call ours. Ours call your tools.

auto-AGENTS™ speaks the Model Context Protocol natively in both directions. Expose your own tools and the engine will reason over them; register ours and your agents gain payments, CRM, ticketing and data access, every call grounded, governed and replayable in a full trace.

2.1Bidirectional MCPcall out to your tools, expose ours to your agents
2.2170+ connectorsCRM, payments, ticketing, data, telephony
2.3Full run tracesreasoning, tool call, result & timing per step
Streaming inference · low latency

Tokens, events and audio, as they happen.

Stream model output token-by-token over SSE or WebSocket and emit structured events for every reasoning step and tool call. Sub-500ms first token keeps voice natural; webhooks deliver run completions, scores and escalations to your systems the instant they occur.

3.1Token & event streamingSSE / WebSocket · partial + final
3.2Webhooks & eventsrun.completed · score.ready · escalation
3.3First-class SDKsTypeScript, Python, Go · typed & async
Deploy anywhere · enterprise security

Run it in your cloud, your VPC, or on-prem.

Model-agnostic and portable, bring GPT, Gemini, Claude or your own on-prem Llama. Pin to a region for data residency, isolate in your VPC, or deploy fully on-prem. Enterprise auth via SSO / SAML / OIDC, scoped API keys, and an immutable audit trail on every call.

4.1Cloud · VPC · on-premregion pinning & data residency
4.2Model-agnostic runtimeyour model or ours, swappable per run
4.3Enterprise auth & auditSSO / SAML / OIDC · scoped keys · immutable log
What teams ship

An engine your builders can stand on.

0

Connectors, day one

Reach across CRM, payments, ticketing, data and telephony over MCP without writing a single integration adapter.

0

First-token latency

Streaming inference keeps voice and chat natural, with structured events for every reasoning and tool step.

0

Concurrent sessions

A horizontally-scaled runtime with admission control, proven at high-concurrency voice volumes.

Trust & compliance

Certified for regulated builders.

Deterministic guardrails before and after every model call, PII vaulting, scoped credentials and your choice of region, with an immutable audit trail behind every API request.

SOC 2HIPAAISO 27001PCIGDPR
FAQ

Questions from developers.

Every capability of auto-AGENTS™, running agents, registering tools, indexing knowledge, voice, evals and analytics, is exposed as a REST + streaming API with no required UI. You can run the entire engine from your own application, your IVR, or your own agents. The dashboard is just one optional client of the same API you get.

auto-AGENTS™ speaks the Model Context Protocol natively. You can expose your own tools as MCP servers and the engine will reason over and call them within your guardrails; you can also register our 170+ connectors so your agents gain payments, CRM, ticketing, data and telephony access. Every call appears in a full run trace, reasoning, tool call, result and timing.

Stream model output token-by-token over SSE or WebSocket, plus structured events for every reasoning step and tool call. First-token latency runs under 500ms to keep voice natural. For asynchronous work, webhooks deliver run completions, QA scores and escalations to your endpoints the instant they happen.

Yes. Bring GPT, Gemini, Claude or your own on-prem Llama, swappable per run with fallbacks. Deploy in our cloud, your private VPC, or fully on-prem, and pin to a region for data residency. SDKs ship for TypeScript, Python and Go.

Enterprise auth via SSO / SAML / OIDC, scoped and rotating API keys, field-level encryption and tenant-scoped key management. Deterministic guardrails run before and after every model call, PII is vaulted, and every API request is recorded in an immutable audit trail. The platform is certified to SOC 2, HIPAA, ISO 27001, PCI and GDPR.

“We had an agent calling our own MCP tools and streaming tokens back into our app in an afternoon. The API is the product, no UI to fight, full traces on every run, and bidirectional MCP meant our existing tools just worked.”
Staff Engineer
Platform team · enterprise fintech

Built directly on the headless engine, agents, tools and knowledge driven entirely over REST + streaming, with every reasoning step and tool call replayable from the run trace.

<1 day
to first streamed run
<500ms
first-token latency
MCP connectors

170+ tools your agents can call on day one.

Salesforce ServiceNow Genesys Amazon Connect Twilio Zendesk Microsoft Dynamics Webex Snowflake Slack Microsoft Teams

Build agents on the headless engine.

Every capability is an API. Bring your stack and your models, we'll wire the engine in.

Built for the people who ship.

Clean APIs, SDKs and webhooks to embed auto-AGENTS™ anywhere, from first call to production in days.

customer enjoying a call
Build with our APIRead the docs ↗
customer smiling on the phone
customer on a call
customer on a call