Repeated tool-call loops occur when an agent cannot recognize progress, failure, or completion and therefore keeps selecting the same action.
A self-hosted agent may repeatedly search the same folder, rerun one command, reopen the same file, or submit an identical API request even though the result cannot change. The visible loop is only the symptom. Its root can sit in the model’s plan, the tool schema, the returned observation, stored agent state, an external retry wrapper, or a missing termination rule. Distinguishing those layers matters because a longer context window or higher turn limit can extend the loop without explaining it.
The Loop Signature Is Repeated Action Without New State
A legitimate agent may call the same tool several times with different arguments or after receiving new evidence. A pathological loop repeats an equivalent call while the task state, available evidence, and predicted next step remain materially unchanged.
LangGraph documents a graph recursion limit for workflows that execute too many steps before reaching a stop condition, including graphs with unintended cycles.
The distinguishing observation is not the raw number of calls. It is whether each call creates a new fact, changes a resource, narrows the plan, or moves the graph toward a terminal state.
Ambiguous Tool Results Leave the Agent Unsure Whether Anything Happened
A tool can return an empty string, generic success message, partial payload, stale cached value, or human-readable error that does not clearly identify success, retryability, or permanent failure.
The Model Context Protocol separates tool execution errors by placing explicit error state in the tool result. When a runtime flattens every outcome into ordinary text, the model must infer whether another call could help.
A loop caused here usually shows the same tool and arguments following an observation that contains no stable completion marker. The tool may be functioning correctly while its response contract remains too vague for the agent to update its plan.
State Writes Can Succeed Outside the Agent but Fail Inside Its Memory
A file may be created, a database row may be updated, or a service may restart while the agent’s stored state still says the action is pending. The next reasoning turn therefore repeats a completed operation.
ReAct-style agents interleave reasoning, action, and observation so that observations update the action plan. If an observation is dropped, assigned to the wrong tool-call ID, truncated, or excluded from the next prompt, the control loop loses the evidence needed to move forward.
This root cause is distinguishable from model confusion because the external system shows progress while the trace presented to the model does not. Replaying only the model turn with the correct observation often produces a different next action.
Retry Layers Can Turn One Failure Into Several Identical Calls
The model may request one tool call while the orchestration framework, HTTP client, queue worker, or task runner retries it multiple times. The final trace can look like agent indecision even when repetition happened below the model layer.
Tenacity’s retry controls separate the decision to retry from stop conditions and retry predicates. A broad retry rule can repeat deterministic validation errors, permission failures, or malformed arguments that cannot succeed without changed input.
Look for identical request IDs, timestamps, exception classes, and model-turn counts. Several executions inside one agent turn indicate runtime retries; one execution per new reasoning turn points back toward agent planning or state interpretation.
Weak Completion Criteria Keep Returning Control to the Tool Router
An agent can finish the requested side effect but lack a machine-readable condition that says the overall task is complete. The router sees another tool-capable model message and sends control back into the action node.
The OpenAI Agents SDK exposes a maximum-turn boundary that raises an exception when a run exceeds its configured turn count.
A turn cap bounds damage but does not identify the root cause. If the trace shows successful tool output followed by another equivalent call, the missing element is usually a completion transition, final-answer route, or state field that the router actually checks.
Tool Descriptions Can Encourage the Same Choice After Every Failure
Overlapping tools, underspecified failure behavior, and descriptions that emphasize capability without limits can make one tool appear optimal on every turn.
CrewAI documents iteration and retry limits as separate agent controls, reflecting the difference between repeated reasoning cycles and repeated execution attempts.
This cause is most visible when arguments vary slightly but the chosen tool never changes, even after the observation proves that the tool lacks access, scope, or required data. The loop is then a selection-policy problem rather than a transport retry.
Context Loss Can Erase the Evidence That a Call Already Failed
Long traces, large tool schemas, verbose outputs, and local-model context limits can push earlier failure details or completion markers outside the effective prompt.
The agent then sees the original task and current tool list but not the observation that ruled out its preferred action. It reconstructs the same plan from incomplete history and appears to forget its own attempt.
ZimaSpace’s article on self-hosted automation agents provides the adjacent boundary: adding more tools expands capability, but reliable orchestration still depends on compact state, explicit outcomes, and bounded execution.
FAQ
Is every repeated tool call an infinite loop?
No. Pagination, polling, chunked processing, and iterative search can legitimately reuse one tool. The key test is whether the arguments, evidence, or task state change between calls.
Does increasing the maximum number of turns solve the problem?
No. It can allow a valid long workflow to finish, but it also gives a zero-progress loop more time to repeat. The trace still needs a verifiable progress or termination condition.
Can a stronger model eliminate tool-call loops?
It may interpret ambiguous observations better, but it cannot recover state that was never returned, distinguish hidden runtime retries, or enforce a stop condition absent from the workflow.
Tech & AI HUB
More to Read

What Causes an AI Agent Planner to Repeat Steps It Already Completed?
Trace repeated planner steps through state persistence, completion evidence, tool-result parsing, context retention, retries, replanning, and stop conditions.

What Causes Permission Errors Only Inside AI Agent Subprocesses?
Compare parent and child identity, filesystem view, environment, capabilities, security policy, and executable path to diagnose subprocess-only denial.

What Causes CPU Saturation When Hardware Transcoding and Video AI Run Together?
Trace CPU saturation across codec offload, pixel conversion, frame copies, AI preprocessing, audio, subtitles, storage, and process scheduling.

