Token Lifetime and Rotation Policies for Agentic Workflows
AI agents need credential rules built for continuous operation, not human login sessions.

Token lifetime and rotation rules built for human logins do not hold up when the thing logging in is an AI agent, and that mismatch is the subject of this piece. The angle is straightforward: lifetime, rotation, and scope have to be treated as controls for the whole workflow, not settings tied to a single account.
Most identity systems in use today were built around a simple picture: a person logs in, does some work, and logs out. OAuth and the credential frameworks built on top of it assume a bounded session with a human somewhere in the loop, making judgment calls about what to click and when to stop. An AI agent doesn't fit that picture, and the attempt to force it into one of two existing boxes creates a real problem. Treating the agent like a service account gives it indefinite standing permissions with no built-in way to limit those permissions to a specific task or a specific window of time. Treating it like a human user instead makes the system assume a kind of judgment about appropriate scope that a model simply does not have.
Four things about agents make this mismatch worse, not better. They run around the clock, so a credential held by an always-on agent racks up exposure over hours or days, not the few minutes a human session usually lasts. Their behavior comes from model reasoning rather than a fixed code path, so a permission set that looked fine on paper can turn out either too loose or too easy to break, depending on what the model decides to do next. And agents spin up and shut down sub-agents at container speed, fast enough that handling their credentials by hand simply isn't possible; automation has to do that work.
Machine identities already outnumber human ones by a wide margin inside most companies, and agentic systems push that ratio even further. The credential surface tied to agents is the main attack surface any serious security plan now has to account for, not a side issue to clean up later.
Token lifetime and blast radius
How long a credential stays valid sets the ceiling on the damage it can do, before any other security control gets a chance to matter. A short lifetime works as prevention. Everything else, scope limits, monitoring, revocation hooks, mostly functions as cleanup after the fact.
Think about the sequence an attacker has to complete after stealing a credential: find it, figure out what it's good for, use it, and then try to move from that one foothold to something bigger. Each of those steps takes time. A token that expires quickly can break that chain partway through, simply by going dead before the attacker's tools finish the first pass.
Leaked secrets get found fast now. Automated bots scan public code commits, paste sites, and CI logs, and they test any exposed key within minutes of it showing up. A long-lived token handed to one of these bots is a running head start measured in hours or days. A short-lived one often expires before the first real exploit attempt even finishes.
The risk compounds in any setup where many agents share a similar authentication pattern. A hundred agents all authenticating with static keys form a mesh where one leaked credential can be walked across the entire network, hopping from system to system as long as the key stays valid. Short lifetimes put a hard cap on how far that walk can go before the credential dies on its own.
A self-contained bearer token can't be revoked the instant it's stolen. The only real lever is waiting for it to expire. That makes lifetime the de facto revocation policy for access tokens in most real deployments, whether or not anyone designed it that way on purpose.
Access-token lifetime for agents: the 5-to-15-minute range
A 5-to-15-minute access-token lifetime is a solid default for most agentic workflows. Going beyond that window calls for compensating controls, and going past one hour without those controls is simply not defensible.
Three tiers cover most production cases. Agents that call only a handful of APIs per session can run on the aggressive end of the range, a few minutes at most, since refresh latency only becomes noticeable on call paths that block while waiting. A moderate lifetime, closer to the middle of the range, works as a common default, balancing the size of the theft window against how often the agent has to refresh. Longer lifetimes only make sense when paired with sender-constrained tokens, using mTLS or DPoP, because a token bound to a specific key is worthless to anyone who can't produce that key. Past one hour without that kind of binding, scope limits and anomaly detection alone can't make up for the size of the window an attacker gets to work in.
The usual objection, that short lifetimes mean constant refresh overhead and slower agents, doesn't hold up well in practice. Refreshing in the background a minute or so before expiry means the agent's actual calls never have to wait on it. The overhead stays invisible at the application layer.
Even platform defaults show the tension here. AWS STS sets a 15-minute minimum lifetime for AssumeRole credentials and a one-hour default, a floor that already is at the aggressive edge of the recommended range for agents making a lot of calls. A "short-lived" cloud credential by AWS's own definition may still be longer than ideal for some workloads. GitHub has moved in the other direction, replacing long-lived personal access tokens with an ephemeral GITHUB_TOKEN scoped to each individual Actions job. That shift is a signal from a major platform that credentials should be scoped to the workflow run, not handed out at the application level and left to sit.
Refresh-token rotation: why per-use rotation is the correct default for agents
Short access tokens only work as a control if the refresh token behind them doesn't quietly become the long-lived secret the whole policy was built to avoid. That's why per-use rotation, not reuse, should be the default for any agent.
The mechanism, laid out in OAuth 2.1 (draft-ietf-oauth-v2-1) §4.3.1, is simple to describe: the agent presents refresh token RT1, the server issues a new access token AT2 along with a new refresh token RT2, and RT1 is invalidated immediately. Each rotation shrinks the replay window down to exactly one use. Binding that refresh token to an mTLS certificate or a DPoP key adds another layer on top: even if someone steals the token during its brief one-use window, it's useless without the matching private key.
One operational detail trips up a lot of agent deployments: when several workers in the same system detect an expired credential at the same time, more than one of them may try to refresh it simultaneously. That race condition can trigger false-positive reuse detection and knock out the entire token family by mistake. The fix is to designate a single process, or use a distributed lock, so only one worker ever presents a given refresh token.
Rotation also creates a failure mode that appears far from the code that caused it. If a refresh token gets revoked on the server side, whether from a security event, a policy change, or an administrator stepping in, the agent's framework gets a 401 it has no way to recover from on its own. Someone has to reauthorize the integration by hand. That fact belongs in incident runbooks, not buried somewhere in code comments.
If an agent's refresh token gets its absolute lifetime extended every time it rotates, the session never really ends. That turns the rotation mechanism into theater. The absolute maximum lifetime needs to be anchored to when the token was first issued, and rotation events should never push that ceiling back.
Scope at the tool level, not the agent level, as a containment boundary
Lifetime and rotation handle the time dimension of credential risk. Scope handles the other one: how much a stolen credential can actually do. Issue a token against an agent's entire possible range of actions, rather than the one tool it's calling right now, and a single leak exposes everything that agent could ever touch, not just whatever task it happened to be running at the moment of theft.
The right unit of scope is the specific tool call in front of the agent at that moment, not the agent as a whole. A token authorized for one tool call shouldn't work if replayed against a different tool, even within the same agent session.
Because an agent's actions come out of model reasoning rather than a fixed sequence of steps, broad permission sets are genuinely dangerous in a way that fixed software isn't. Scoping wide and planning to tighten later means living with an overprivileged credential for however long it takes to get around to fixing it, and that window is usually longer than anyone intends. Current practice is moving toward treating credentials as artifacts tied to a single workflow run rather than to an application's whole lifetime: issue on demand for one task, revoke when that task ends, and never carry the same credential forward into an unrelated task.
The tightest version of this idea shows where the field is headed. The Agent Identity Protocol (AIP) proposal introduces Invocation-Bound Capability Tokens, which tie identity, a narrowed set of permissions, and a record of where authority came from into one append-only chain built for each individual invocation. It runs in a compact JWT mode for single-hop calls and a chained Biscuit-token mode for cases where authority passes through multiple agents. Pairing narrow scope with a short lifetime multiplies the benefit of each: a token that's both narrow and short-lived limits how much an exploit can reach and how long it has to work.
Multi-agent delegation chains and single-agent token policy limits
A sub-agent that inherits the full authority of the orchestrator that spawned it is the concrete failure this section is about. Everything so far has dealt with a single agent holding a single credential. Once agents start delegating work to other agents, a new surface opens up, and the policies built for one agent don't automatically cover a chain of them.
In a standard OAuth delegation flow, the token handed down carries the original scope unchanged through every tool call and every handoff to a sub-agent. There's no built-in OAuth mechanism for the delegating agent to narrow that scope at the point where it hands off work. A sub-agent spun up to do one narrow task can end up holding the same authority as the agent that created it.
Some proposals point toward a fix without yet being one. Google DeepMind's "Intelligent AI Delegation" paper proposes Delegation Capability Tokens, built on Macaroons or Biscuits, designed around the idea that scope should narrow by default at each delegation step. The paper lays out that philosophy but hasn't shipped a working protocol or implementation.
Revocation gets harder inside a chain, too. If the primary agent loses access, sub-agents that already received delegated, offline-attenuated tokens may keep working, since there's no automatic path that propagates revocation to tokens already handed further down the chain. A security analysis of agentic communication protocols found a related gap in ACP: because its enforcement of short-lived tokens and JWS timestamps is optional rather than required, replay attacks stay viable in longer sessions that skip the timestamp check. The protocol supports the right behavior. It's only as strong as whatever a given deployment chooses to turn on.
The Coalition for Secure AI's 2026 guidance on delegation tokens lays out a clearer standard: keep delegation tokens short-lived, on the order of seconds to minutes, tightly scoped, and revocable through a centralized service capable of near-real-time invalidation. Runtime components should subscribe to revocation events as they happen, rather than checking on a polling schedule that leaves a gap between revocation and enforcement.
Revocation as a real-time operational control, not a post-incident cleanup step
An audit log reviewed after an incident and a revocation triggered after data has already left the building are both just postmortems. Revocation only works as a safety control if it fires faster than the agent's next tool call, not sometime after the damage is done.
A self-contained bearer access token can't be revoked the instant it's stolen. The only option is to wait for it to expire. Short lifetime has to be the primary control, with explicit revocation serving as the backstop behind it.
Revoking the refresh token is the practical lever available in the moment. Doing so stops the agent from getting new access tokens going forward, but the current access token stays valid until it naturally expires. Short access-token lifetime and refresh-token revocation each cover the other's gap, so the two controls have to work together. When a refresh token gets revoked mid-incident, every downstream system still holding access tokens minted from it keeps operating normally until those tokens run out on their own. Response plans need to build that grace period into their assumptions rather than treating revocation as an instant kill switch.
Rotation overlap windows need careful calibration, too. If the window is too tight, legitimate agents retrying a failed refresh will trip false-positive family revocation. If it's too loose, that same overlap becomes a window an attacker can exploit through replay. An agent that can act on its own shouldn't hold a credential that outlives the action it was authorized to complete. Revocation hooks should fire automatically the moment a task finishes, not only when something has already gone wrong.
Token caching in agent runtimes: where correct policy is most frequently undone in practice
Correct lifetime, rotation, and scope policy gets undone constantly at the caching layer, and that's where most real-world token failures in deployed agent systems actually come from. The policy on paper is usually fine. What breaks it is a runtime that caches credentials in a way that quietly extends their effective lifetime or blurs together scopes the policy was designed to keep separate.
The cache key has to be the combination of user, agent, and audience together, not any one of those alone. When caching spans users, one person's credential can end up served to a different user's request. When caching spans agents, the scope boundaries that tool-level issuance was built to enforce collapse into one shared pool. When caching spans audiences, credentials meant for one resource server leak into requests aimed at a completely different one.
Proactive refresh is the pattern that keeps this from happening in practice: fetch a new token a set margin of time before the old one expires, on a background thread, well before any foreground call needs it. That approach keeps foreground requests from ever blocking on an expired token, and it also heads off the race condition from earlier, where several processes fire refresh attempts at once and trip false-positive reuse detection. Policy only holds up if the caching layer respects the same boundaries the policy was written to draw.


