Agent Reliability

observabilityopentelemetrytelemetrystandardstracing

Observability standards for agents (OpenTelemetry GenAI)

The emerging standard layer for agent telemetry: OpenTelemetry's GenAI semantic conventions define spans, metrics and events for model calls, tool execution, MCP interactions and agent operations, so traces from different stacks become comparable. Still in active development — adopt it for vocabulary and interoperability, pin versions, and expect churn until stabilization.

Why this wins its question: Tells the adopter what the standard actually covers today and what its in-development status means operationally (pin versions, expect churn) — most references either ignore the conventions or oversell their maturity.

Claims

Every assertion below is bound to registered sources and carries its own confidence. Weight them; do not treat the page as uniformly authoritative.

  1. OpenTelemetry maintains dedicated semantic conventions for generative AI — spans, metrics and events for GenAI clients, MCP and provider-specific operations — covering agent spans, tool execution and token usage.

    confidence 0.9OpenTelemetry semantic conventions for generative AI · primary

  2. The GenAI conventions live in their own repository under Apache-2.0 and are in active development, with the schema URL still marked TODO — the vocabulary is adoptable, the stability guarantees are not yet.

    confidence 0.85OpenTelemetry semantic conventions for generative AI · primary

  3. Framework backing for instrumenting deployed AI: NIST AI RMF makes Measure a core function across the lifecycle, which post-deployment telemetry operationalizes.

    confidence 0.85NIST AI Risk Management Framework (AI RMF 1.0, NIST AI 100-1) · primary