Build Flows

AI agents & MCP · October 10, 2026 · 14 min read

What to Log on Every Agent Tool Call: An MCP Telemetry Blueprint

A field-level specification for MCP tool call telemetry: five event types, common and per-event fields, privacy rules that keep raw payloads out, an OpenTelemetry span structure, and derived metrics including argument auto-correction and hallucinated tool names.

By Charley Forey, founder of Build Flows

The short answer: Emit five structured, versioned events: mcp.call.started, mcp.call.completed and mcp.call.failed for every tool call, plus mcp.session.summary when a session ends and mcp.chain.summary when a run of tool calls inside one model response finishes. Every event carries a request ID, a session ID, a hashed user ID, the environment and trace IDs. Never put raw tool inputs, raw outputs, prompts, tokens or personal data in this stream; log hashes, sizes, schema identifiers and flags instead. Trace each model inference as a parent span, each tool call as a child, and each external API call nested under that. Then compute reliability, product, cost and quality metrics downstream, including two signals most teams miss: arguments the server had to auto-correct, and tool names the model made up.

It is based on the telemetry specification we drafted for MCP tool calls during our agent work at Trimble, rewritten as a vendor-neutral blueprint. It is a companion to our article on the Tool Runtime MCP gateway, which shows a working gateway and its per-call log. That article is about the product. This one is the field-level spec: what to emit, what to keep out, and what to compute from it.

Two records, two jobs

Teams often try to make one log do everything. It works better as two:

Call log (debug and audit)Telemetry stream (this blueprint)
PurposeAnswer "why did the agent say that?" for one callRun SLOs, product analytics, cost attribution and quality tuning across all calls
ContentsInput and output, where policy allows, plus who and through which keyMetadata, hashes, sizes, timings, flags. No raw payloads
AudienceA small group with access to the underlying dataInfra, platform, product and data teams, dashboards, alerting tools
RetentionShort, or per your audit requirementsRaw events for weeks; aggregates for years

The Tool Runtime gateway captures input and output on purpose, because that is what makes a single bad answer debuggable. The telemetry stream goes to more people and more tools, so it has to be safe to share. Keep them separate, join them on request_id, and you get both.

Goals and non-goals

Write these down before you add a single field. Ours:

Goals

  • Enforce reliability and SLOs for tool calls.
  • Trace model and tool interactions end to end for debugging.
  • Measure tool adoption and usage.
  • Attribute cost by tool, session and tenant.
  • Support audit and compliance reviews.

Non-goals

  • Full payload logging of tool inputs and outputs.
  • Logging raw personal data or secrets.
  • Reconstructing whole conversations from telemetry alone.

The non-goals do as much work as the goals. They are what you point to when someone asks to "just add the prompt" to the event.

Event taxonomy

EventWhen it firesWhy it exists
mcp.call.startedA tool call is initiatedCaptures intent: which tool, from where, with what schema, before anything can fail
mcp.call.completedA tool call returns successfullyLatency breakdown, size, retries, tokens and estimated cost
mcp.call.failedA tool call fails terminallyError category, retryability and timing
mcp.session.summaryA session ends or times out from inactivityOne row per session for product and cost analysis
mcp.chain.summaryA sequence of tool calls within one model response cycle completesShows how tools are combined, how deep chains go, and where they are abandoned

Every started event should end in exactly one completed or failed event with the same request_id. A started event with no terminal event is itself a signal: a crash, a timeout the server never saw, or a client that disconnected.

Timeline of one session: in the first model response a call starts and completes and another starts and fails, then chain.summary fires; in the second response a call starts with no terminal event; session.summary fires at session end or timeout. A legend lists the five events and when each fires.Where the five events fire in a session, and the orphaned start that should raise a flag.

All events must be structured JSON, versioned (event_version), forward compatible (consumers ignore fields they do not know) and aligned with OpenTelemetry where possible.

Common fields on every event

FieldRequiredNotes
event_nameYesFor example mcp.call.started
event_versionYesFor example 1.0. Bump on breaking changes only
timestampYesISO 8601, UTC
request_idYesUnique per tool invocation. The join key to the call log
session_idYesUnique per user session. Hash it if the raw value works as a credential
user_id_hashYesKeyed hash, never the raw ID (see privacy rules below)
environmentYesprod, staging or dev
tenant_idOptionalThe company or organization in a multi-tenant deployment
app_idOptionalWhich integration or agent made the call
trace_id, span_idOptionalFor OpenTelemetry correlation. Make them required as soon as tracing is live

Call event fields

On mcp.call.started:

FieldRequiredNotes
mcp_server_nameYesThe logical server, not the hostname
tool_nameYesThe tool identifier as the model called it
invocation_sourceYesmodel, system or user
tool_versionOptionalSchema version of the tool
tool_schema_hashOptionalHash of the tool's input schema. Detects schema drift between what the model saw and what the server runs
model_name, model_versionOptionalWhen a model initiated the call
input_schema_validOptionalResult of pre-validation
parallel_group_idOptionalShared by calls the model issued concurrently

On mcp.call.completed:

FieldRequiredNotes
latency_msYesTotal call duration
successYesAlways true on this event; keeps queries simple
network_latency_ms, tool_execution_ms, model_reasoning_msOptionalThe breakdown that tells you where time went
output_size_bytesOptionalLarge outputs cost context and money
retry_count, fallback_triggeredOptionalHidden instability shows up here before it shows up as failures
tokens_input, tokens_output, estimated_cost_usdOptionalFor cost attribution

On mcp.call.failed:

FieldRequiredNotes
error_typeYesOne of validation, timeout, auth, rate_limit, internal
retryableYesWhether a retry could succeed
error_codeOptionalService-specific code, such as the stable codes in your error envelope
error_message_hashOptionalHash, not the raw message, which can contain data
latency_ms, retry_countOptional

Keep error_type to a small closed set. Five categories are enough to drive alerting. The detail belongs in error_code, which should match the stable error codes your server already returns (we describe that envelope in hardening MCP servers for production).

A completed event in practice looks like this:

{
  "event_name": "mcp.call.completed",
  "event_version": "1.0",
  "timestamp": "2026-10-10T14:03:22.418Z",
  "request_id": "req_7f3c",
  "session_id": "sess_hash_91ab",
  "user_id_hash": "uh_4d20",
  "tenant_id": "tenant_a",
  "environment": "prod",
  "trace_id": "4bf92f3577b34da6a3ce929d0e0e4736",
  "span_id": "00f067aa0ba902b7",
  "mcp_server_name": "erp-ap",
  "tool_name": "list_review_queue",
  "latency_ms": 1840,
  "tool_execution_ms": 1610,
  "output_size_bytes": 18233,
  "retry_count": 1,
  "success": true,
  "argument_auto_corrected": false,
  "tool_chain_position": 2
}

Notice what is missing: no invoice numbers, no vendor names, no amounts, no query text.

Session and chain summaries

mcp.session.summary rolls a session into one row: total_tool_calls, successful_calls, failed_calls, session_duration_ms and unique_tools_used are required; token totals, total estimated cost, max_tool_depth and fallback_used are optional. Emit it on explicit session end and on inactivity timeout, or you will undercount every session that simply went quiet.

mcp.chain.summary covers one model response cycle: how many tool calls it made, in what order, how deep the nesting went, whether the chain completed, and how long it took. This is where you see that a user question about a job's committed cost took five calls when it should take two, or that a chain keeps dying at the same step. If you route requests through a gateway tool or capability layer, record which tool or capability each request resolved to as part of the chain; it is the data you need to tune tool routing.

Model and tool interaction signals

These fields are where telemetry stops being generic API monitoring and starts improving the agent. They feed tool schema design, prompt changes and model routing decisions.

FieldMeaningWhere to detect it
tool_requiredThe model indicated the tool call was mandatoryFrom the model request, if your client exposes it
tool_hallucinatedThe model called a tool that is not registeredThe server or gateway, at dispatch. Record the attempted name in a hashed or allowlisted form
argument_auto_correctedThe system repaired the model's arguments before running the callYour validation or normalization layer, for example coercing a single item into a list or fixing a date format
tool_response_usedThe model actually referenced the tool's output in its answerHard to detect reliably; treat as a sampled, offline judgment
tool_chain_positionOrder of this call within the chainThe client or orchestrator
chain_depthNesting level when tools call toolsThe orchestrator

Two of these deserve extra attention.

Argument auto-correction is a quiet tax. Every repair your server makes is a place where the tool's schema or description failed to tell the model what it wanted. If one tool's correction rate climbs, fix the schema: tighter types, clearer field descriptions, examples. Our servers normalize inputs before validating them (single items and lists both become one shape), and that kind of quiet repair is exactly what you should count, because it hides which tools confuse the model.

Hallucinated tool names tell you the model's picture of the toolset is wrong. Common causes: a tool was renamed, a description promises something another tool does, or a large catalog is being exposed where a gateway would serve better. The approach we describe in the MCP gateway pattern for construction APIs cuts this sharply, because the model picks from a handful of tools instead of hundreds.

Add your own workflow outcomes alongside these. If your server returns structured responses such as need_more_info, approval_required and policy_blocked, count them per tool. Repeated need_more_info on the same field means the agent's instructions are missing something, and approvals requested versus granted is the most honest measure of whether an agent is ready for more autonomy.

Privacy and compliance rules

Never log in the telemetry stream:

  • Raw tool inputs.
  • Raw tool outputs.
  • Secrets, API keys and access or refresh tokens.
  • Full prompts, unless gated and redacted under a specific policy.
  • Personal data, unless encrypted and handled under your compliance rules.

Allowed:

  • Payload hashes.
  • Schema identifiers and schema hashes.
  • Size metrics.
  • Boolean flags.
  • Classification labels, such as "contains financial data".

Privacy boundary: raw tool inputs, full prompts, and secrets and tokens are dropped by a typed emitter with a field allowlist; raw outputs pass only as size and hash, user IDs as a keyed HMAC hash, upstream error text as an error code plus message hash, and the input schema as a schema hash.What crosses the emitter: sizes, hashes and codes, never payloads, prompts or secrets.

Three implementation details make the rules hold up:

  • Use a keyed hash for user IDs. Email addresses and employee numbers are low-entropy. A plain SHA-256 of one can be reversed by hashing a list of likely values. Use an HMAC with a secret key held outside the telemetry pipeline, so analysts can count and join on user_id_hash without being able to recover the person.
  • Hash error messages, keep error codes. Raw error text from upstream systems often echoes the request, including record values. The code tells you the category; the hash lets you group identical messages without storing them.
  • Enforce it in code, not in a wiki. Build events with a typed emitter that only accepts approved fields, and add a test that fails if a new field is not on the allowlist. A redaction step at the collector is a good second line, not the first.

Retention: keep high-cardinality raw events for 30 to 90 days, aggregated metrics for 12 to 24 months, and audit logs for as long as your compliance requirements say.

For construction companies, the sensitive fields are often less obvious than names: invoice amounts, vendor bank details, payroll lines, bid numbers, and the contents of contracts and change orders. None of that belongs in a shared telemetry stream.

OpenTelemetry span structure

Map the events onto traces so one request can be followed from the model to the upstream API:

agent.turn                              (root: one user request)
└── model.inference                     (parent: the model call that chose tools)
    ├── mcp.tool_call  list_review_queue   (child: one span per tool call)
    │   ├── http.client  GET /invoices      (nested: external API call)
    │   └── http.client  GET /invoices?page=2
    └── mcp.tool_call  get_review_packet
        └── http.client  GET /invoices/{id}

Trace waterfall: agent.turn contains model.inference, which parents two tool-call spans, list_review_queue with two nested HTTP calls to the invoices endpoint and get_review_packet with one nested HTTP call. Trace context propagates from client to MCP server to upstream API.The span tree as a waterfall: inference, tool calls and nested API calls on one trace.

Guidelines:

  • One span per tool call, carrying request_id, mcp_server_name, tool_name, the outcome and the interaction flags as attributes. The events above can be emitted as span events or as separate log records that share trace_id and span_id.
  • Propagate W3C trace context (traceparent) from the client to the MCP server and from the server to the upstream API, so a slow ERP call shows up as a slow nested span rather than a mystery.
  • Parallel calls are siblings. Calls in the same parallel_group_id share a parent and overlap in time; the flame graph shows whether parallelism is actually saving time.
  • For streaming responses, one span per stream with events for progress, not a span per chunk.
  • Keep high-cardinality values out of metric labels. User, session and request IDs belong on spans and logs, not on metrics.

OpenTelemetry also has semantic conventions for generative AI, still marked experimental at the time of writing, that include tool-call attributes. Map your field names onto them where they line up, and keep your own names where they do not.

Derived metrics

None of these are logged directly. Compute them downstream from the events.

Map from event sources to metric families: call events feed reliability, cost and quality; session summaries feed product and cost; chain summaries feed product. A callout says to start by joining cost with success rate.Four metric families computed downstream from three event sources.

FamilyMetricBuilt from
ReliabilityTool success rate per tool and per tenantcompleted / (completed + failed)
Reliabilityp50, p95 and p99 latencylatency_ms on completed events
ReliabilityError rate by category; timeout frequencyerror_type on failed events
ProductShare of sessions that invoke any toolsession summaries
ProductAverage tools per session; tool adoption ratesession summaries, started events
ProductTool chain frequency; abandoned chainschain summaries
CostCost per session, per tool, per tenantestimated_cost_usd, token counts
CostHigh-cost, low-success toolscost joined with success rate
QualitySchema validation failure rateinput_schema_valid
QualityArgument correction rateargument_auto_corrected
QualityHallucinated tool ratetool_hallucinated
QualityRetry rate and fallback frequency per toolretry_count, fallback_triggered

The cost-and-quality join is the one that pays for the whole exercise. A tool that is expensive and fails often is your first fix. A tool that is cheap and never used is a candidate for removal, which also shrinks the catalog the model has to reason over.

Dashboards and alerts

Three dashboards cover most needs:

  • Reliability: success rate heatmap by tool, p95 latency by tool, error category breakdown, timeout trend.
  • Product: tool usage over time, top tools, average tools per session, chain visualization.
  • Cost: cost per tenant, cost per tool, cost per successful outcome, high-cost failure analysis.

Add a fourth, smaller quality panel set for the agent team: auto-correction rate, hallucinated tool names, need_more_info by field, and approvals requested versus granted.

Alerts to start with:

  • Error rate or p95 latency on a tool crosses a written threshold. On our Vista AP server, the rollback triggers for configuration canaries were an error rate above 5% or p95 latency above 4 seconds; set yours from a baseline.
  • A burst of auth failures, which usually means an expired secret or a misconfigured audience rather than an attack, but deserves a look either way.
  • A sustained rise in rate_limit failures against one upstream system, which often precedes the integration user being throttled or locked out.
  • Started events without terminal events, above a small baseline.
  • Any new hallucinated tool name, or a jump in one tool's argument correction rate after a deploy. These are tickets, not pages.

Rolling it out

  1. Write the goals, non-goals and field allowlist and get infra, product and data to agree on them.
  2. Emit the three call events from the gateway or server, with the common fields and a typed emitter.
  3. Turn on tracing and make trace_id and span_id required.
  4. Add session and chain summaries once call events are trustworthy.
  5. Add interaction signals, starting with hallucinated tools and auto-correction, because they are cheap to detect at the server.
  6. Build the dashboards, then the alerts, then review quality metrics weekly with whoever owns the agent's prompts and tool schemas.

If you run several MCP servers, emit from the gateway so every server gets the same events for free, then add server-specific detail inside the tool span.

Checklist

  • Goals and non-goals written; non-goals include payload logging.
  • Five event types defined, versioned and forward compatible.
  • Every started event ends in exactly one completed or failed event.
  • Common fields on every event, with keyed hashing for user IDs.
  • No raw inputs, outputs, prompts, tokens or personal data in the stream, enforced by an allowlist test.
  • Error types are a closed set; error messages are hashed.
  • Traces run from model inference to tool call to external API, with context propagated.
  • Reliability, product, cost and quality metrics computed downstream.
  • Auto-correction and hallucinated tool rates reviewed after every tool or prompt change.
  • Retention set for raw events, aggregates and audit logs.

Where to go next

Frequently asked questions

What should I log on every MCP tool call?

A started event and exactly one completed or failed event per call, each with a request ID, session ID, hashed user ID, environment, server and tool name, and trace IDs. Completed events add latency, size, retries and cost; failed events add a closed error type and whether the failure is retryable.

Should agent telemetry include tool inputs and outputs?

Not in the shared telemetry stream. Keep raw payloads in a separate, access-controlled call log for debugging and audit, and join the two on request ID. The telemetry stream should carry only hashes, sizes, schema identifiers and flags.

How do I structure OpenTelemetry spans for AI agents?

Make the model inference the parent span, each tool call a child span, and each external API call a nested span under the tool call. Propagate W3C trace context from the client through the MCP server to the upstream API, and keep user and session IDs on spans rather than metric labels.

What is a hallucinated tool call?

A call to a tool name that is not registered on the server. Count them at dispatch. A rising rate usually means a tool was renamed, descriptions overlap, or the model is choosing from too large a catalog.

How long should MCP telemetry be retained?

A common split is 30 to 90 days for high-cardinality raw events, 12 to 24 months for aggregated metrics, and audit logs for as long as your compliance requirements specify.

Next step

Have a workflow in mind?

Start with a readiness review: the task, the data and tools it needs, the access boundaries, and how a pilot would be evaluated.

Prefer email? charley@buildflows.ai

Get the next guide in your inbox

Field Notes: practical guides and new walkthroughs, about once a month.

Field Notes

Practical guides and new walkthroughs on construction data and automation, roughly monthly.

Keep learning