Trigger.dev chat.agent: Durable AI Chat With No Timeouts

Jeff Liu··4 min read·AI Agents
Trigger.dev chat.agent: Durable AI Chat With No Timeouts
ListenTrigger.dev chat.agent: Durable AI Chat With No Timeouts
0:00
--:--

Key Takeaways

  1. 1Trigger.dev announced chat.agent on August 10, 2026. It has been generally available since July 2 and running in production since June.
  2. 2chat.agent has served millions of sessions and more than 84 years of compute, restoring conversations after a crash.
  3. 3Graham Tremper of Arena said giving every conversation a real machine made durable agents more straightforward to build, and that default tracing made debugging sessions easy.
  4. 4Head Start roughly halved both time to first token and total turn length in Trigger.dev's own tests.

Trigger.dev announced chat.agent on August 10, 2026, a way to build durable AI chat experiences that run on a dedicated Linux machine with no timeouts and keep streaming through refreshes and crashes. Writing it up, CEO Matt Aitken noted the feature has been generally available since July 2 and running in production since June, arriving having already served millions of sessions and more than 84 years of compute.

Traditional chat backends struggle with the stateless nature of request/response cycles, forcing developers to write everything to a database, reach for Redis to get durable streams, and hand slow work to a background worker they then have to coordinate. chat.agent replaces that with a stateful, durable environment for AI agents.

What is chat.agent's core function?

chat.agent gives each conversation its own Linux machine, which starts when the first message arrives, runs as long as the work takes, and holds state in memory between turns. The machine sleeps when nobody is typing and wakes where it left off. You write the turn as a Trigger.dev task that takes messages and returns a stream, then point the AI SDK's useChat at it, and the API route in between goes away.

Because a turn has no timeout, slow tool chains and sub-agents are just work rather than something to chop into request-sized pieces. Trigger.dev reports that in production one run in twenty lasts longer than 36 minutes, well past where a normal request would have been cut off.

Arena built its Agent Mode on chat.agent and runs it in production at scale.

Every conversation gets a real machine, which made our durable agents much more straightforward to build. The default tracing and observability make viewing and debugging agentic sessions incredibly easy.

— Graham Tremper, Arena

How does chat.agent improve AI agent durability?

Memory carries across a sleep, so a variable you set on turn three is still there on turn twenty tomorrow and an expensive lookup you already did stays done. The stream is durable too: refresh mid-response and it replays from where the browser stopped reading, without re-running the model.

A crash is where the distinction matters. If the machine dies you get a new one and the conversation comes back with it, because the conversation is written down. The heap is not. In-memory variables start empty again, which is the same rule as any long-running server. Trigger.dev's own guidance is to put anything you cannot afford to lose in a database and keep memory for speed. What changes here is that memory now lasts the whole conversation instead of a single request.

Deploys do not interrupt anything either. A run stays on the version it started on, and moving a conversation onto new code is an explicit call.

Furthermore, chat.agent offers Head Start, which runs the first Large Language Model call inside your own warm server while the agent boots in parallel, then hands the conversation over mid-turn. In Trigger.dev's tests it roughly halved both time to first token and the length of the whole turn.

For operations requiring approval, a tool with no execute function ends the turn with the call still open. The agent suspends, the person takes as long as they take, and their answer resumes the run. Because it is suspended it is not billed, so an approval can sit overnight or over a weekend.

The platform integrates with existing AI SDK implementations, using streamText on the server and useChat on the client. Only the new message goes over the wire, since history accumulates on the server. Long conversations stay affordable through compaction and prompt caching.

Every turn is a span in the dashboard, and an AI metrics view ships with every project covering spend, calls, time to first chunk, tokens per second, latency percentiles by model and cost by task and provider, with no instrumentation to write.

One detail that gets less attention than it deserves: Trigger.dev is Apache 2.0 and chat.agent is part of it. The implementation can be read and self-hosted, and a chat agent is billed as compute time only while it is actually running.

Feature

Traditional Chat Backends

Trigger.dev chat.agent

Compute Environment

Stateless API endpoint with timeouts

Dedicated Linux machine per conversation

State Management

External store required for every turn (Redis, database)

In-memory state persists across turns and sleep; a database is still required for anything that must survive a crash

Conversation Durability

Limited by request timeouts, requires manual retry logic

No turn timeouts; conversations resume after days or a crash

First Turn Speed

Agent boots on every request, potentially slow

Head Start runs the first call on a warm server; roughly halved TTFT in Trigger.dev's tests

Cost Efficiency

Billing for active compute regardless of user input

No charge while suspended and waiting on a person

Observability

Requires manual instrumentation for metrics

Built-in tracing and an AI metrics dashboard for cost, tokens, latency

License

Varies; often proprietary

Apache 2.0, self-hostable

Related Articles

More insights on trending topics and technology

The Signal

What shipped in AI this week, with the sources.

One email a week.