Cosmic

MCP Agents All Logged Out at Once? How to Handle Clerk Auth Errors on Your MCP Server

Saad Ahmed · Founder5 min read

Short answer

As of October 2026, an MCP server that verifies Clerk OAuth tokens must separate Clerk's verdicts on a token from errors about its own server credentials. Refuse an agent with 401 and error="invalid_token" only when Clerk says the token is inactive, revoked, expired or unknown (404). Treat a verify 401 or 403, rate limits, 5xx and timeouts as Clerk being unavailable: answer 503 with Retry-After and keep already-confirmed agents working for a bounded grace period, or one Clerk-side error logs out every agent at once.

Key takeaways

  • A 401 from Clerk's token verify endpoint is about your server's secret key, not the agent's token; Clerk's own SDK reads it as InvalidSecretKey.
  • Refuse agents only on verdicts about the token itself: active false, revoked, expired or 404.
  • Answer Clerk-side failures with 503 and Retry-After, no sign-in challenge, and a bounded grace for agents you already confirmed.
  • Keep your own revocations instant during an outage, and log every failed Clerk call with its status and error code.
  • If your MCP server never pushes messages, answer GET with 405; an idle stream is something a proxy can cut every 15 minutes.

If you're building an MCP server with Clerk as the OAuth provider, so people can bring their own agents to your app, this post is for you. It covers the one error-handling decision that decides whether a hiccup at Clerk logs out a single agent or every agent at once, plus a second bug that quietly dropped our agents' connections every 15 minutes.

Written in October 2026. Clerk's API, the MCP specification and client behavior are described as they stood that month.

We hit both in production on ecco, an agent work coordinator where people and their AI agents share one task board. Agents such as Claude Code and Codex connect to it over MCP, and Clerk issues their OAuth tokens.

"When MCP is table stakes for building apps in 2026 and BYO-agent is the approach users want in their agentic app experiences, you can't risk brittle, silently failing auth for your app." Saad Ahmed

What happened

At 19:51 UTC on Saturday, October 3, our server started answering every Clerk-verified agent with 401, credential invalid. Claude Code sessions, a claude.ai connection and every other OAuth agent were refused within the same minute.

None of the usual causes applied:

  • The tokens were valid. One still had about 21 hours left, and daily token refreshes had worked all week.
  • Nothing shipped. There was no deploy between October 2 and October 5.
  • Nothing changed in our data. No Clerk webhooks arrived that day, and no agent credential was created, revoked or expired.
  • The person's Clerk account was fine. It wasn't locked or banned, and its password and OAuth apps were unchanged.

The failure also looked random. Our server remembers a successful Clerk check for up to an hour to save round trips, so some sessions kept working until their remembered check ran out at 20:58, while others failed at once. Recovery took a manual sign-in in every client, about 30 hours later.

Saad's reaction when he saw it: "We can't have that happen with real people and real users."

Why every agent failed together

Our server checks each agent token with two calls to Clerk: a POST to `oauth_applications/access_tokens/verify`, authenticated with our Clerk secret key, and a GET to `/oauth/userinfo` with the agent's token.

The bug was in how we read Clerk's errors. Any 400, 401, 403 or 404 from those calls became "this agent's token is invalid". But a verify 401 isn't about the agent's token. Clerk returns 401 when it rejects the server's own secret key, and that answer comes back the same for every token you send. So one Clerk-side rejection of our server's call turned into a mass logout.

Clerk's own SDK reads these codes the same way. In `@clerk/backend`, the error handler for token verification maps a 401 to `InvalidSecretKey` and only a 404 to `TokenInvalid`.

What Clerk's verify responses mean, and how to answer the agent
Clerk responseWhat it's aboutAnswer the agent with
200 with the token activeValid tokenServe the request
200 with active false, revoked or expiredThis token401 with error="invalid_token"
404This token (Clerk doesn't know it)401 with error="invalid_token"
401 (clerk_key_invalid)Your server's secret key503 with Retry-After
403 (authorization_invalid)Your server's access to Clerk503 with Retry-After
429, 5xx, timeout, unreadable bodyClerk's availability503 with Retry-After
400Ambiguous: a malformed token or request401 for a token you never confirmed; treat as unavailable for one you did

We couldn't confirm which status Clerk returned at 19:51, because our server didn't log it then. A server-key rejection fits every symptom. The fix below holds whatever the code was, and now every failed Clerk call is logged with its status.

How to fix it, step by step

  1. Sort Clerk's answers into two groups: verdicts on the token (active false, revoked, expired, 404) and failures about your server or Clerk's availability (401, 402, 403, 422, 429, 5xx, timeouts, malformed bodies).
  2. Refuse the agent with 401 only for verdicts on the token.
  3. Answer everything else with 503 and a Retry-After header, and send no sign-in challenge with it, so clients wait and retry instead of starting a new login.
  4. Keep agents you've already confirmed working through a short outage. Ours gets up to one extra hour past its normal one-hour remembered check, never past the token's own expiry.
  5. Keep your own revocations instant during an outage. Removing a person from a workspace, revoking an agent or disabling a workspace still refuses on the next call, because those checks live in your database, not at Clerk.
  6. Log every failed Clerk call with the call, the status, Clerk's error code and your verdict. Never log tokens, keys or Clerk's message text.
  7. Replay the incident in a test: make the fake Clerk reject the server key, and assert that confirmed agents keep working, new tokens get 503, and genuine revocation still refuses at once.

The trade-off is small and worth stating. A revocation made only at Clerk, during an outage, with its webhook lost, takes effect up to one hour late. Everything revoked on your side still takes effect immediately.

Send a sign-in signal clients can act on

When a token really is bad, say so in the standard way. A 401 that carries `WWW-Authenticate: Bearer error="invalid_token"` (RFC 6750, section 3), plus the `resource_metadata` URL for your protected resource (RFC 9728), tells MCP clients to re-authenticate. Without the error code, some clients retry the same token and end in a "needs-auth" state.

Keep the two answers apart: a 401 with `invalid_token` means "sign in again", while a 503 with no challenge means "wait and retry". Mixing them up sends clients into a login loop during an outage.

The second bug: a connection that dropped every 15 minutes

While reading the client logs, we found another pattern. Claude Code's long-lived GET connection to our MCP endpoint ended in an error exactly every 15 minutes, around the clock, 99 times in 24 hours. After three in a row, the client tore down its transport and reconnected. It recovered every time, so nothing in the system flagged it, but the churn buried real failures in noise.

The exact interval points at a proxy limit on long-lived requests. The real fix was simpler than tuning it: our server never sends anything on that stream.

The MCP specification (2025-11-25, Streamable HTTP, "Listening for Messages from the Server") says a server must either return an event stream for that GET or answer 405 Method Not Allowed. Our server is stateless, advertises no resource subscriptions and no list-change notifications, and never sends unsolicited messages, so we now return 405 with `Allow: POST`. Since that change went live, Claude Code's logs show no dropped streams, and tool calls work as before.

Know before your users do

A fix only helps if the next failure is visible. We added three things:

  • An alert for the team when three or more different agents are refused in five minutes. One person's expired token is normal; several agents at once means something upstream broke.
  • A notice for the person whose agent was refused, in the app's Inbox, written for a human: which agent stopped and how to reconnect it. One notice per agent, closed when it signs in again.
  • A soak test. An unattended agent runs a daily check for seven days, with no one touching it, and counts sign-in prompts and dropped streams.

Check these before you ship an MCP server on Clerk

MCP server auth checklist
CheckWhy
A verify 401 or 403 becomes 503, never "invalid token"Those codes are about your server, so treating them as the user's fault logs everyone out
401 responses carry error="invalid_token" and resource_metadataClients know to sign in again instead of retrying a dead token
503 responses carry Retry-After and no Bearer challengeClients wait out an outage instead of starting a login
Confirmed agents get a bounded grace during an outageA short Clerk incident doesn't become a user-facing outage
Your own revocations skip the graceRemoving someone works immediately, even when Clerk is down
Every failed Clerk call logs status and error codeThe next incident is diagnosable from your logs alone
GET to your MCP endpoint returns 405 if you never push messagesNo idle stream for a proxy to cut
An alert fires when several agents fail at onceYou hear about it before your users do

"Even though I love Clerk out of the box, like with any infra tooling for app development, you need to be able to bend it to your will if needed. And then build custom scaffolding around the standard tooling to make it work if it can't bend the way you like. In this instance, agentic auth is a new space and documentation is limited, so here's hoping you found this writeup useful." Saad Ahmed

FAQ

Does a 401 from Clerk's token verify endpoint mean the user's token is invalid?

No. As of October 2026, Clerk returns 401 (clerk_key_invalid) when it rejects your server's secret key, and Clerk's own SDK maps a verify 401 to InvalidSecretKey. Only a 200 with the token inactive, revoked or expired, or a 404, is a verdict on the token itself.

Why does my MCP client say needs-auth after a 401?

The server refused the token and the client couldn't refresh it, so it waits for a manual sign-in. Send WWW-Authenticate: Bearer error="invalid_token" with your resource_metadata URL so clients know to re-authenticate, and don't send a 401 at all when the real problem is your server or the identity provider.

How long should an MCP server keep agents working during a Clerk outage?

Long enough to ride out a short incident, with a hard cap. Ours keeps an already-confirmed agent working up to one hour past its normal one-hour remembered check, never past the token's own expiry, while revocations made in our own database still take effect on the next call.

Should an MCP server support the GET stream in Streamable HTTP?

Only if it sends messages the client didn't ask for. The MCP specification (2025-11-25) lets a server answer GET with 405 Method Not Allowed when it offers no stream. A stateless server with no subscriptions or list-change notifications has nothing to send, and an idle stream invites proxy timeouts.

Can Clerk be the OAuth provider for an MCP server?

Yes. As of October 2026, Clerk can act as the OAuth authorization server for an MCP server, issuing scoped tokens to agents such as Claude Code and Codex that your server verifies on each request.

Sources

  1. Clerk docs: Verify OAuth tokens with Clerk
  2. Clerk docs: Backend API errors
  3. Clerk on GitHub: @clerk/backend token verification (handleClerkAPIError)
  4. Clerk docs: Build an MCP server with Clerk
  5. MCP specification 2025-11-25: Streamable HTTP, listening for messages from the server
  6. MCP specification 2025-11-25: Authorization
  7. RFC 6750: Bearer token usage, WWW-Authenticate errors
  8. RFC 9728: OAuth 2.0 Protected Resource Metadata
Saad Ahmed

Founder of Cosmic. Nearly 15 years in strategy, management and sales as a digital strategist, director of sales and product manager. He closed over $25M in new revenue at VMware Pivotal Labs, launched digital products at Capital One, and helped Viget grow from 30 to 75 people. He started Cosmic and built the Revenue Design® process after advising friends whose companies struggled to close deals and keep revenue steady.