AI Agents

OpenAI Dots: Your AI Agent in Production — It Won’t Work

JG

Jared H. Garr

CEO, Rebirth Distribution

OpenAI Dots: Your AI Agent in Production — It Won’t Work

Reading time: 6 min

Key Takeaways

  • Persistent agents like OpenAI’s Dots need deterministic infrastructure — not cloud black boxes with API rate limits.
  • Production failures kill trust: network blips, context drift, credential management, and silent errors that never make the logs.
  • Boring infrastructure beats flashy agents: start with queues, gates, and review, then earn trust incrementally.

OpenAI just announced Dots, a persistent agentic assistant powered by their new GPT-6 Astra model. At DevDay, they described them as “remarkably capable, always-on agents built to handle everything.” Floating cartoons. Bubbly personas. The promise that your digital buddies will work on your behalf, continuously, in the background.

The demo worked. Production didn’t. Here’s why.

I run Rebirth Distribution. We design and deploy production-grade agent systems for startups — n8n pipelines, VPS infrastructure, OpenClaw/Hermes orchestration. Not demo-grade. The kind that has to hold when real customers are pinging your endpoints at 3am.

I’ve seen what happens when you take a persistent agent out of the sandbox and drop it into a real environment. It’s not pretty.

What Actually Happens in Production

Let me be specific. You deploy a Dots agent to “monitor customer feedback, implement bug fixes, and ship requested features.” Here’s the sequence of failures nobody talks about:

  • API rate limit hit — your cloud provider throttles the agent’s LLM calls. It stops mid-task. No alert. No retry. Just silence.
  • Network partition — the agent loses connection for 30 seconds during a deploy window. It retries the same operation three times. Now you have three database entries where you wanted one.
  • Context drift — the “always-on” agent has been running for 72 hours. Its memory window has cycled through thousands of messages. It forgets the original goal and invents a new one.
  • Credential bloat — you gave it access to your GitHub. Then Slack. Then Jira. Then your production vault. Now you have a single compromised bot that owns your entire stack.
  • Silent hallucinations — the agent writes a bug fix that looks correct, runs without errors, and deletes the wrong column. No exception. No traceback. Just data loss.

That’s not automation — that’s a liability. The real cost is: downtime hours, data corruption, team dependency on a system no one can debug.

Most people get this wrong. They evaluate agents based on what they can do. The correct evaluation is based on what they do when things fail.

Why the “Always-On” Model Breaks

OpenAI is selling an ideal: Dots that pursue “user-defined goals continuously in the background with minimal oversight.” From an architecture standpoint, that’s not a feature. That’s a guarantee of cascade failure.

Any system that runs continuously needs state management, idempotency, deterministic retry, and structured observability. LLM-based agents give you none of that out of the box. They give you non-deterministic logic inside a loop with a finite context window.

This isn’t theory. I’ve deployed agents on n8n workflows that were just 20 steps long — none of them were LLM-based in the control path. Pure deterministic runners. We still had to add three layers of failure handling before they stopped paging me at night.

Now imagine a Dots agent with the same intent, but instead of predefined workflow steps, it uses GPT-6 Astra to decide what to do next. That’s a state machine where every transition is a coin flip.

The Specialist Illusion

OpenAI envisions “specialist Dots” with provisioned identities, credentials, and tools. This is where anyone with production experience should be alarmed.

The real cost is credential sprawl. Each Dots instance needs access tokens. Each token is a blast radius. A “customer support Dot” with Slack access can leak your private channels if its context gets prompt-injected. A “developer Dot” with GitHub push rights can merge to main. Do you trust six independent LLM agents to manage your infrastructure?

There’s a reason we lock services down with short-lived tokens, scoped IAM roles, and mandatory review gates. An agentic assistant that “acts independently” and “runs in the background” doesn’t fit that model. It fits an idealized model where you ignore every operational principle we’ve spent a decade developing.

OpenAI is integrating with Microsoft Agent 365 security controls — good. That tells me they know this is a problem. But security controls don’t fix architectural fragility. They just contain the blast radius after the failure.

What Actually Works Right Now

If you’re a startup looking at Dots and wondering whether to skip building your own stack — the answer is: no. Not yet. Not for anything that touches production.

Here’s what actually holds up:

  • Deterministic first, LLM last. Build the workflow in n8n, Zapier, or custom orchestration with explicit steps. Put the LLM only where judgment is genuinely required.
  • Human-in-the-loop gates. Any write action — code, data, message — goes through review until the agent has earned trust over thousands of invocations.
  • Queue-based invocation. Never run agents continuously. Run them on-demand with a message queue and a dead-letter queue. When they fail, they go somewhere you can inspect.
  • Scoped credentials. One agent, one service, one token, one TTL. Revoked daily. Rotated automatically.
  • Observability before autonomy. If you can’t trace every agent decision from input to output, you don’t deploy it. Period.

This is where OpenClaw and Hermes fit into our stack — they’re the orchestration layer we built because cloud providers don’t give you the primitives you need for deterministic agent control. Not glamorous. Not cartoonish. But they hold.

The Bottom Line

Dots might be great inside ChatGPT. It might save you time scheduling meetings, drafting emails, or summarizing notes. That’s fine. That’s consumer-grade.

But if you’re running a startup and thinking of replacing any piece of your production automation with Dots — wait. Watch how it fails for the people who go first. Then look at the failure patterns. If you can accept those costs and contain them, fine. If you can’t, build the deterministic layer yourself.

Keep the pipeline boring. Ship the product exciting.

← Back to Latest