Repository navigation
fix(agent-sessions): span-kind classification - #1126
JeremyFunk wants to merge 4 commits into
Conversation
…ention one Symptom: a span with no known operation was a model call when its name contained "chat" or "completion", a tool call when it contained "tool", and an agent step when it contained "agent" or "workflow". LiteLLM proxy's `auth /chat/completions` span doubled LLM calls (list: 6 for 3 requests); LlamaIndex's `BaseWorkflowAgent.call_tool` and `aggregate_tool_results` counted every tool call 3x; DSPy's `ChatAdapter.__call__`, LangGraph's `tools` node, Spring AI's `chat_client` advisor and an OpenAI Agents workflow named "... chat" were counted as calls; Haystack's `haystack.agent.step.llm` read as an agent step, so raw Haystack sessions showed 0 LLM calls. Cause: `classifyAiSpan` (packages/agent-sessions/src/session-turns.ts) and its SQL transcription `genAiIsLlmCallCond` / `genAiIsToolCallCond` (packages/domain/src/tinybird/gen-ai-columns.ts) fell back to substring matches on the span name. Fix: past the operation, a span is a tool call by its tool name, a model call by its model, and otherwise by an exact framework span name (`AI_INFERENCE_SPAN_NAMES` / `AI_TOOL_SPAN_NAMES` in @maple/domain/gen-ai): `haystack.agent.step.llm`, `haystack.agent.step.tool` and Effect AI's `Toolkit.handle`, the only calls in the replayed captures that carry no operation, model or tool name. Everything else is an agent step. The session summary query gains the same names. Migration 0035 recreates ai_trace_index_mv with the new flags (forward-only, nothing backfilled); local store v25 -> v26. Seen in: trace-capture docs_litellm_proxy, llamaindex_agents, dspy_agents, langgraph_agents, spring_ai_agents, openai_agents_sdk_agents, haystack_agents, effect_ai_agents.
…call counts Symptom: on a LiteLLM proxy session the session page counted 9 LLM calls where the list counted 6 (3 real requests; the other 3 are the proxy's `auth` spans, fixed separately). The proxy's FastAPI server span `POST /chat/completions` carries only `gen_ai.request.model` and http.*, so the ingest gateway does not stamp it and the list's index never holds it, but the page classified it as an inference span. Cause: `classifyAiSpan` (packages/agent-sessions/src/session-turns.ts) applied its model and name fallbacks to any span with a decoded gen_ai key (`isAiSpan`), stamped or not. The session summary query (`summaryMeasures_` in packages/query-engine-integrations/src/ai/ai-sessions.ts) did the same for its model fallback. Fix: past the operation, a span the gateway did not stamp is "other" — the app's own work — on the page, and the summary's model fallback requires the vendor stamp, as its tool fallback already did. The list is unchanged: it only ever saw stamped spans. Seen in: trace-capture docs_litellm_proxy (EU org service docs-verify-litellm, proxy session).
|
Warning Review limit reachedNext included review available in 16 minutes. View limit detailsLimit details: You’ve used all 4 included reviews currently available. You've used all free OSS reviews for now. Wait for the free limit to reset to keep reviewing this public repository. Review configuration: ⚙️ Run configurationConfiguration used: defaults Review profile: CHILL Plan: Advanced Run ID: ⛔ Files ignored due to path filters (2)
📒 Files selected for processing (22)
Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out. Comment |
Maple reviewConfidence 4/5 · likely safe to merge Span-kind classification now reads a model or tool call from the span's tool name, its model, or an exact framework span name instead of substrings, and drops spans the ingest gateway never stamped. The page, the
What was checked
|
There was a problem hiding this comment.
Devin Review found 1 potential issue.
1 flag not posted on this PR by your GitHub settings — view it in Devin Review. (Configure)
| const isToolCall = operation.in_(...AI_TOOL_OPERATIONS).or( | ||
| operation | ||
| .eq("") | ||
| .and(isAi) | ||
| .and(toolName.neq("").or(model.eq("").and($.SpanName.in_(...AI_TOOL_SPAN_NAMES)))), | ||
| ) |
There was a problem hiding this comment.
🟡 Unknown-operation tool calls vanish from summaries
A stamped tool span with an unknown nonempty operation and a toolName contributes no summary tool call. The index classifier accepts unknown operations, so the list counts that call.
Learn more
The summary query aggregates spans for the session detail page. Its tool-call fallback currently runs only when the operation is empty. The page classifier and index classifier instead fall through whenever an operation is not one of the known operations. The summary therefore disagrees with both on a tool span carrying an open-set operation.
Example: A stamped span with gen_ai.operation.name=invoke_tool and gen_ai.tool.name=search counts as one tool call on the list and page but zero in the summary.
Recommended fix: Gate the summary fallback on the operation being outside all known operation sets, as genAiIsToolCallCond does. Preserve the vendor check and tool-name-before-model precedence.
| const isToolCall = operation.in_(...AI_TOOL_OPERATIONS).or( | |
| operation | |
| .eq("") | |
| .and(isAi) | |
| .and(toolName.neq("").or(model.eq("").and($.SpanName.in_(...AI_TOOL_SPAN_NAMES)))), | |
| ) | |
| const isToolCall = operation.in_(...AI_TOOL_OPERATIONS).or( | |
| operation | |
| .notIn(...AI_INFERENCE_OPERATIONS, ...AI_RETRIEVAL_OPERATIONS, ...AI_TOOL_OPERATIONS, ...AI_AGENT_OPERATIONS) | |
| .and(isAi) | |
| .and(toolName.neq("").or(model.eq("").and($.SpanName.in_(...AI_TOOL_SPAN_NAMES)))), | |
| ) |
Was this helpful? React with 👍 or 👎 to provide feedback.
There was a problem hiding this comment.
Fixed in 84d061c: the summary's tool fallback now uses the same not-in-known-operations guard as its model-call fallback, the index and classifyAiSpan.
…nown operation (#5) Symptom: the session summary query counted a tool call off its tool name only when the span had no operation at all, so a stamped span with an operation the convention does not name (LangSmith `chain`, Spring AI `framework`) and a tool name was a tool call on the page and in the index but not in the summary. Cause: `summaryMeasures_` (packages/query-engine-integrations/src/ai/ai-sessions.ts) guarded its tool fallback with `operation = ''`, where its model-call fallback, `genAiIsToolCallCond` and `classifyAiSpan` all fall back past any operation outside the known sets. Fix: the tool fallback uses the same not-in-known-operations guard. Found by the independent verification of the #5 commit.
Maple reviewConfidence 4/5 · likely safe to merge Aligns the session summary's model- and tool-call rules with the page's
FindingsNote · F1 · New model-call test passes off the tool-call rule's SQLtests · The expected substring stops at What was checked
Copy all findings (1)
|
…an-classification # Conflicts: # packages/query-engine-integrations/src/ai/ai-sessions.ts
|
Note Maple is reviewing this pull request at |
Span-kind classification fixes for Agent Sessions: which spans count as model calls and tool calls, on the list (
ai_trace_index), the session summary query, and the session page (classifyAiSpan). Found while replaying the agent-tracing guide captures (bugs #5 and #37 of the verification run).#5: calls read off span names that merely mention one (6a1e021)
Symptom. A span with no known operation was a model call if its name contained
chatorcompletion, a tool call if it containedtool, and an agent step if it containedagentorworkflow. From the replayed captures:auth /chat/completions(scopelitellm, no attributes): list shows 6 LLM calls for 3 requests.BaseWorkflowAgent.call_tool+aggregate_tool_results: every tool call counted 3x on list and page (4x with HITL).ChatAdapter.__call__(13 per run, beside 13 realLM.__call__), LangGraphtools/tool_executor/ChatPromptTemplate, Spring AIspring_ai chat_client/message_chat_memory/ Tool Calling Advisor, an OpenAI Agents workflow named "… chat": counted as calls.haystack.agent.step.llm(no operation, no model): read as an agent step, so raw Haystack sessions showed 0 LLM calls.Chat.generateTextwrapsLanguageModel.generateText, which already carriesgen_ai.operation.name=chat: each call counted twice.Cause.
classifyAiSpan(packages/agent-sessions/src/session-turns.ts:77-85before this PR) and its SQL copygenAiIsLlmCallCond/genAiIsToolCallCond(packages/domain/src/tinybird/gen-ai-columns.ts:184-231, vianameLooks) fell back to substring matches on the lower-cased span name.Fix. Past the operation, a span is a tool call by its tool name, a model call by its model, and otherwise only by an exact framework span name:
AI_INFERENCE_SPAN_NAMES = ["haystack.agent.step.llm"],AI_TOOL_SPAN_NAMES = ["haystack.agent.step.tool", "Toolkit.handle"](packages/domain/src/gen-ai.ts). Everything else is an agent step. I checked every vendor inapps/ingest/src/ai_session.rsagainst the captures for spans that carry no operation, model or tool name. These three names are the only real calls among them:gen_ai.operation.nameat ingest. Its phase spans are unstamped.ai.model.id/ai.toolCall.name.call_llm, LiteLLM SDKlitellm_request, Semantic Kernelchat.completions: carry a model.The same rule is used in three places: the page (
classifyAiSpan), the MV (genAiIsLlmCallCond/genAiIsToolCallCond) and the session summary (summaryMeasures_inai-sessions.ts, which had no name rules and now reads the same exact names).Migration.
0035_ai_trace_index_call_classificationdrops and recreatesai_trace_index_mv(verbatim emitter DDL). The target table is unchanged,requiredForIngest: false. Forward-only: rows materialized before it keep their oldIsLlmCall/IsToolCalluntil rawtraces' 30-day TTL ages them out. The page and the summary read raw spans, so they are correct immediately. Local store v25 -> v26 (local-0025-to-0026-ai-trace-index-call-classification: drop the view, bootstrap recreates it).Test.
packages/agent-sessions/src/session-turns.test.tscovers the Haystack/Effect names that must still classify and the eight capture names that must not. SQL tests are inai-span-columns.test.tsandai-sessions.test.ts; the migration test is inmigrations/index.test.ts.#37: session page counts the LiteLLM proxy's FastAPI span as a model call (a967430)
Symptom. On a LiteLLM proxy session the page counted 9 LLM calls, the list 6, for 3 real requests. The other 3 on the list are the
authspans that #5 fixes. The proxy's FastAPI server spanPOST /chat/completionscarries onlygen_ai.request.modelandhttp.*. The gateway does not stamp it and the index never holds it, yetinspect_spancalled it "AI agent span — inference".Cause.
classifyAiSpan(session-turns.ts:73) applied its fallbacks to any span with a decoded gen_ai key (isAiSpan,ai-integrations.ts:363), stamped or not. The summary query's model fallback (ai-sessions.ts:1618-1623) also ignored the stamp, although its tool fallback already required it.Fix. Past the operation, a span the gateway did not stamp is
other(the app's own work) on the page. The summary's model fallback now requires the vendor stamp. The list is unchanged because it only ever held stamped spans.Test.
session-turns.test.ts"classifies an unstamped span as other…" (the proxy span's shape fromdocs_litellm_proxy). There is also an SQL assertion inai-sessions.test.ts.Follow-up commit (84d061c). The summary's tool-call fallback required an absent operation (
operation = ''). The page, the index and the summary's own model-call rule all fall back past any operation outside the known sets. It now uses the same not-in-known-operations guard.Correction to the 6a1e021 commit message. The message says the three names are "the only calls in the replayed captures that carry no operation, model or tool name". That holds for the
docs_*captures only. Across all 113 captures, the OpenInference-only DSPy/smolagents spans (covered by #1121) and the native LlamaIndex package (below) also rely on data the page could not read.Decisions and known effects
chat gpt-5) reading. An emitter that writes that name should also writegen_ai.operation.name, and every such span in the captures does.llama-index-observability-otelpackage: known limit. For thellamaindex_agentscapture, the page shows 16 LLM calls / 17 tool calls before this PR, 0 / 0 after, and the true figure is 8 / 3. The package's spans carry no operation, model or tool name; the LLM calls exist only asLLMChatStartEventspan events. The old counts came from substring matches that also countedcall_tool,aggregate_tool_results,_prepare_chat_with_toolsand nestedastream_chatpairs, so no name rule gets close to the truth. The LlamaIndex guide steers users to the OpenInference instrumentor, which classifies correctly.unknown:other): spans named{name}.toolwith nogen_ai.tool.nameused to be tool calls by substring and are now agent steps. There is no capture to verify this against. The principled fix is to translatetraceloop.span.kindthe wayopeninference.span.kindis translated; that's a follow-up.ai-sessions.ts). fix(agent-sessions): session checks, vitals and turn labels #1120 also editssession-turns.ts(turn anchors and labels, notclassifyAiSpan). Since Widget lab + heatmap layout fix #37, an unstamped span with no operation isother, so it can no longer open a turn as an agent root.