40 Retries in 90 Seconds: Claude Code Versus Hermes in Prod

40 Retries in 90 Seconds: Claude Code Versus Hermes in Prod

What our circuit breaker caught at 3am looked like a runaway retry loop, triggered by a single transient 529 hitting a retry loop with no backoff. Most coverage of Claude Code treats it as a finished coding agent. Anthropic's own production guidance says otherwise: retry with backoff, circuit breakers, and budget limits are things you build yourself, not things you get by default. That gap between SDK and service is where this incident lives, and it's also where Claude Code and Hermes Agent turn out to be solving different problems entirely. This post answers which of the two problems you actually have, and what it costs to solve either one properly.


Here's the comparison: Claude Code, which Anthropic ships as an SDK you assemble into your own service, versus Hermes Agent, a gateway style tool that comes with more of the operational scaffolding pre built, including a cron scheduler that runs jobs out of a local directory. Neither is finished in the sense of needing zero glue code. Claude Code fits teams who want control over failure modes and are willing to write the reliability layer themselves. Hermes fits teams who want scheduled, semi autonomous agent runs without writing a scheduler from scratch. People compare these two constantly, but they're not solving the same problem.


Claude Code Puts Reliability Entirely On The Integrator

Retry With Backoff: Correct Failure Handling Flow

Retry With Backoff: Correct Failure Handling Flow

1. API Call Fails
→
2. Check Status Code
↓
3a. Non retryable (400) → Raise Immediately
or
3b. Retryable (429, 500, 502, 503, 529)
↓
4. Exponential Delay + Jitter, Retry (max 5)
↓
5. Success or Raise After Max Attempts

Without this logic, a single transient 529 with no backoff becomes a runaway retry loop, the exact incident described in the article.

Source: Source: Article code logic, "40 Retries in 90 Seconds"


Claude Code ships as an SDK, so the five areas that matter in production, infrastructure, security, reliability, observability, and performance, are things you assemble rather than things you get by default. Anthropic's production deployment guidance covers infrastructure concerns like secrets management, rate limiting, multi region failover, and monitoring; security concerns like permission modes, tool allow and deny lists, sandbox isolation, PII filtering, and audit logging; reliability concerns like circuit breakers, retry with backoff, graceful degradation, and budget limits; observability concerns like structured logging with correlation IDs, cost tracking, and health checks; and performance concerns like token counting, prompt caching, and concurrent tool limits. That's a wide surface area for something marketed as a coding agent.


A working prototype and a production deployment of Claude Code can be separated by a few hundred lines of harness code that never touch the model itself. I've built the retry with backoff piece three separate times for three separate clients, and the version that stuck was not the naive exponential backoff most tutorials show. It distinguished between retryable errors, like 429 and 529, and non retryable ones, like a 400 from a malformed tool call, because retrying a malformed request burns rate limit budget without ever succeeding. The incident that opened this post is the concrete case: a 529 with no backoff logic behind it turned one transient error into a runaway loop. That's the exact failure this code prevents.


import time
import random

RETRYABLE_STATUS = {429, 500, 502, 503, 529}

def call_with_backoff(fn, max_attempts=5):
    for attempt in range(max_attempts):
        try:
            return fn()
        except APIStatusError as e:
            if e.status_code not in RETRYABLE_STATUS:
                raise  # do not retry malformed requests
            if attempt == max_attempts, 1:
                raise
            delay = min(60, (2 ** attempt) + random.uniform(0, 1))
            time.sleep(delay)

The permission mode system is the part people underestimate until it bites them. Claude Code supports tool allow and deny lists at a granularity that lets you permit Read and Grep while blocking Bash in a given deployment context, which matters when the same agent harness serves both a trusted internal engineering team and a customer facing support flow. Sandbox isolation on top of that keeps a prompt injected instruction from becoming a filesystem write in someone else's tenant. Get the permission mode wrong once in a multi tenant setup and a security review will find it before any stack trace does. Five domains, zero of them optional. That's what the checklist actually amounts to.


That checklist is also the dividing line from Hermes. Every item above is something Claude Code expects the integrator to build. Hermes takes a different position on one of those five domains, reliability of scheduled execution, by shipping part of it directly.


Hermes Ships A Cron Scheduler That Claude Code Lacks By Design

Claude Code vs Hermes Agent: What You Get by Default

Claude Code vs Hermes Agent: What You Get by Default

Domain Claude Code (SDK) Hermes Agent (Gateway)
Infrastructure Build yourself Partially prebuilt
Security Build yourself Partially prebuilt
Reliability Build yourself Partially prebuilt
Observability Build yourself Partially prebuilt
Scheduling Build yourself Cron scheduler included

Both tools require glue code; Hermes ships more operational scaffolding pre built, while Claude Code leaves reliability decisions entirely to the integrator.

Source: Source: Anthropic production deployment guidance, as described in article


Hermes Agent's most distinctive production feature, based on my own testing of its gateway, is a built in cron scheduler. Jobs live as files in ~/.hermes/cron/, and a background thread inside the gateway process periodically checks that directory for work to execute. Small design decision, large practical consequence: scheduled, unattended agent runs are a first class citizen in Hermes rather than something bolted on with a separate cron entry or a Kubernetes CronJob calling out to an SDK script.


In my own testing, I dropped a job definition into the directory to run a repo health check every morning and the gateway picked it up on a subsequent poll without a restart. That's a genuinely convenient property if the use case is closer to scheduled autonomous maintenance, dependency audits, or nightly log summarization than to interactive pair programming. The tradeoff: the polling based pickup is an eternity if sub second job pickup matters, and the mechanism is a directory watch rather than an event driven trigger, so a job definition written incompletely, say by a text editor that writes in two chunks, could in principle get picked up mid write. I haven't seen that race condition fire in practice, but I wouldn't want to debug it in a job that fires payroll adjustments.


# ~/.hermes/cron/nightly-audit.yaml
name: nightly-dependency-audit
schedule: "0 6   *"
prompt: |
  Check package.json and requirements.txt for outdated
  dependencies with known CVEs. Summarize findings only,
  do not modify files.
permissions:
  tools_allow: [Read, Grep, WebSearch]
  tools_deny: [Bash, Write, Edit]
timeout_seconds: 300

The permission block in that job file does the same conceptual work as Claude Code's allow and deny lists, but scoped per scheduled job rather than per session or per deployment. That's a meaningfully different unit of control. A Claude Code deployment typically applies one permission profile across a service. Hermes lets each cron job carry its own, so a nightly audit job can stay read only while a separate weekly cleanup job gets Write access. The polling interval is the detail that defines what Hermes is for.


Scheduling is only half of running a job unattended. The other half is knowing what happened after it ran, and that's where the two tools diverge again.


Observability Failures Look Different In Each Tool

The 5 Production Domains Claude Code Leaves to You

The 5 Production Domains Claude Code Leaves to You

1

Infrastructure

Secrets, rate limiting, failover, monitoring

2

Security

Permission modes, sandbox isolation, PII filtering

3

Reliability

Circuit breakers, retry with backoff, budget limits

4

Observability

Structured logging, cost tracking, health checks

5

Performance

Token counting, prompt caching, concurrency limits

0 of 5

Optional

Every domain requires custom harness code

Source: Source: Anthropic production deployment guidance, as described in article


Both tools claim observability, but the failure modes I've actually hit in each differ in kind. Claude Code's structured logging with correlation IDs works well when a team follows the documented pattern of tagging each session with a request ID that threads through every tool call, but the practice is opt in. Nothing stops a team from shipping a Claude Code integration with print statements and no correlation ID at all. I inherited exactly that codebase from a team that moved fast and skipped it.


Cost tracking is the other half of observability that gets skipped under deadline pressure, and it generates the angriest Slack message when it's missing. Token counting before a request estimates cost, but the number that matters operationally is cumulative spend against a budget limit. Claude Code's guidance is to implement that budget limit directly, checking accumulated cost against a threshold and triggering graceful degradation, such as falling back to a smaller model or a cached response, when spend approaches the ceiling.


import logging
import uuid

logger = logging.getLogger("agent.session")

def run_session(prompt, budget_usd=5.00, spent_so_far=0.0):
    correlation_id = str(uuid.uuid4())
    if spent_so_far >= budget_usd:
        logger.warning(
            "budget_exceeded",
            extra={"correlation_id": correlation_id, "spent": spent_so_far}
        )
        return fallback_response(prompt)

    logger.info("session_start", extra={"correlation_id": correlation_id})
    result = claude_code_client.query(prompt)
    logger.info(
        "session_complete",
        extra={"correlation_id": correlation_id, "cost": result.cost_usd}
    )
    return result

Hermes, by contrast, provides basic execution logs per cron job out of the box because the gateway process is already the thing running every job, giving it a natural place to log start, finish, and failure per job file. What it doesn't provide for free is cross job correlation. If a nightly audit job and a weekly cleanup job both touch the same repository and something goes subtly wrong across both, tracing that interaction means stitching together two separate job logs by hand. Two halves of the same problem, and neither tool solves the whole thing.


Logging what a session did is one problem. Controlling what a session is allowed to touch, especially once more than one user or one customer shares the same deployment, is a separate problem, and it's where the gap between the two tools becomes structural rather than cosmetic.


Multi Tenant Isolation Is Where Prototypes Actually Break


A prototype that works for one developer on one laptop almost never survives contact with a second tenant, and this is the single most common reason I've seen a Claude Code proof of concept get sent back for rework. The sandbox isolation guidance in Anthropic's own production checklist exists precisely because the default assumption, one user, one working directory, one set of credentials, breaks the moment customer A and customer B hit the same backend service.


The fix in practice: run each tenant's session inside its own sandboxed working directory with its own scoped API key or scoped permission profile, plus PII filtering at the boundary before anything gets logged or sent upstream. I watched a team skip the PII filtering step because it felt like a compliance checkbox rather than an engineering problem. That gap left their correlation IDs leading back to raw user input sitting in plaintext logs, exactly the kind of thing an audit is likely to flag.


import re

EMAIL_RE = re.compile(r"[\w.+-]+@[\w-]+\.[\w.-]+")
SSN_RE = re.compile(r"\b\d{3}-\d{2}-\d{4}\b")

def redact_pii(text: str) -> str:
    text = EMAIL_RE.sub("[REDACTED_EMAIL]", text)
    text = SSN_RE.sub("[REDACTED_SSN]", text)
    return text

def log_session_input(correlation_id, raw_prompt):
    logger.info(
        "session_input",
        extra={"correlation_id": correlation_id, "prompt": redact_pii(raw_prompt)}
    )

Hermes sidesteps a chunk of this problem because its cron jobs are typically scoped to a single operator's home directory and a single set of credentials by design. That makes it less naturally multi tenant out of the box than a Claude Code SDK deployment fronted by a custom API layer. Hermes reads closer to a personal or small team automation tool wearing a gateway's clothes, while Claude Code reads closer to infrastructure that teams are expected to multi tenant themselves. One tenant per gateway instance is the practical ceiling worth planning around for Hermes today.


Isolation determines who can reach a given session. Budget limits determine how much damage an unattended session, tenant isolated or not, can do before anyone notices. That guardrail is the last piece, and it's the one both tools leave to the integrator in different ways.


Budget Limits Fail Silently More Often Than Teams Expect


Budget limits are the reliability feature everyone implements last and regrets not implementing first. I've watched this exact sequence play out at two different companies. A Claude Code integration ships without a hard budget ceiling. A scheduled batch job or a support bot with an unusually chatty user runs longer than expected. The first anyone hears about it is a finance team asking why the Anthropic bill climbed sharply that month.


The graceful degradation pattern that the production checklist recommends, falling back to a smaller model or a cached response near budget, only works if the budget check happens before the expensive call, not after. I made this mistake myself early on: logging cost after the fact showed exactly how much I had overspent, which was informative and useless in the same breath. The fix was moving the check to a pre flight gate.


MODEL_FALLBACK = {
    "claude-opus-4-5": "claude-sonnet-4-5",
    "claude-sonnet-4-5": "claude-haiku-4-5",
}

def select_model(preferred_model, spent_usd, budget_usd):
    ratio = spent_usd / budget_usd if budget_usd else 0
    if ratio < 0.8:
        return preferred_model
    fallback = MODEL_FALLBACK.get(preferred_model)
    if fallback:
        logger.warning("budget_near_limit_fallback", extra={"ratio": ratio})
        return fallback
    raise BudgetExceededError(f"No fallback available, ratio={ratio:.2f}")

Hermes cron jobs need the same discipline but at a different layer. A timeout_seconds field on a job definition caps runtime, but it doesn't cap spend directly, so a job that makes several expensive tool calls within its timeout window can still blow past a reasonable cost expectation before the clock runs out. The pattern worth adopting across both tools: treat budget limit and timeout as two separate guardrails that both need explicit configuration, because neither tool infers one from the other. I wouldn't trust a production agent deployment without an explicit, pre flight budget check written into the code by hand. Not one.


That brings the comparison back to the incident this post opened with. The 3am circuit breaker did its job because someone had written retry logic, backoff, and a budget check by hand, exactly the kind of harness code Claude Code expects an integrator to supply and Hermes would have needed just as much if the runaway loop had been a cron job instead of a live session. Neither tool would have caught that 529 for free. The choice between them isn't about which one is more finished. It's about which gaps you'd rather be responsible for: the five open domains of a Claude Code deployment, or the scheduling convenience and narrower tenancy model of Hermes. Pick based on that, not on which tool's marketing sounds more finished.