DEVELOPER DOCS

SDK

Python SDK

On this page

Installation

Shell
pip install opexia-trace

Optional extras

ExtraInstallPullsUse when
livepip install 'opexia-trace[live]'richYou want the opexia live terminal dashboard
shipcheckpip install 'opexia-trace[shipcheck]'pyyamlYour Ship Check policy is YAML (JSON needs no extra)
litellmpip install 'opexia-trace[litellm]'litellmYou route models through LiteLLM
openllmetrypip install 'opexia-trace[openllmetry]'traceloop-sdkYou already use OpenLLMetry
adapters-langchainpip install 'opexia-trace[adapters-langchain]'langchain-coreYou use LangChain
adapters-microsoftpip install 'opexia-trace[adapters-microsoft]'(none)You use Microsoft agent frameworks

Core dependencies installed always: opentelemetry-api/-sdk/-exporter-otlp (1.x), pydantic 2.x, tiktoken, httpx.


init()

Call once at process startup, before anything you want traced.

Python
import opexia.trace

opexia.trace.init(
    org_id="acme",
    workspace_id="3f2b1c8e-9a4d-4f11-b7c2-1e5a8d3f6b90",
    project_id="checkout-agent",
    backend_url="https://ingest.opexia.dev",
    api_key="opx_live_xxxxxxxxxxxxxxxxxxxx",
    transport="direct",
)

Parameters

All parameters are keyword-only.

NameTypeRequiredDefaultDescription
org_idstr✅—Organization identifier
workspace_idstr✅—Workspace UUID the dashboard reads
project_idstr✅—Logical project/service name; groups traces
backend_urlstr✅—Ingest base URL — https://ingest.opexia.dev
api_keystr✅—Workspace API key (opx_live_…)
collector_endpointstr➖http://localhost:4317OTLP/gRPC collector. Used only when transport="collector"
wal_pathstr➖.opexia-wal/spans.jsonlCrash-recovery write-ahead log
auto_instrumentbool➖TruePatch known LLM client methods automatically
fail_openbool➖FalseIf True, init failures log instead of raising
sampler_ratefloat➖1.0Fraction of traces sampled, 0.0–1.0
transportstr | None➖"collector""collector" or "direct" — see below

Choosing a transport

transport="direct"transport="collector" (default)
PathYour process → ingest.opexia.dev over HTTP/JSONYour process → local collector (gRPC) → Evigauge
Needs a collectorNoYes
Uses collector_endpointNo — ignoredYes
Best forServerless, containers, quick startsExisting OTel infrastructure, local buffering, fan-out

Override without a code change:

Shell
export OPEXIA_TRANSPORT=direct

The explicit transport= argument always wins over the environment variable.

Collector users: your otlphttp exporter must set encoding: json. The protobuf default returns 415 from Evigauge ingest, and it fails silently from the collector's point of view — the single most common cause of "no spans arriving".

YAML
exporters:
  otlphttp:
    endpoint: https://ingest.opexia.dev
    encoding: json                       # REQUIRED
    headers:
      x-opexia-api-key: opx_live_xxxxxxxxxxxxxxxxxxxx

What init() does

  1. Stores config, reachable via get_config().
  2. Fetches the workspace capture_text flag from GET /v1/workspaces/{ws}/sdk-config. Fail-closed — if the lookup fails, capture_text is False and no text bodies are sent.
  3. Builds the exporter for the chosen transport.
  4. Replays any WAL from a previous crashed process, then installs DurableBatchSpanProcessor.
  5. Registers auto-instrumentation when auto_instrument=True.

Steps 3–5 are wrapped by fail_open. set_config() runs first and outside that guard, so get_config() works even in a degraded state.

fail_open

ValueBehaviour on init failure
False (default)Raises. Your process will not start mis-instrumented.
TrueLogs the exception and continues, untraced.

Use fail_open=True in production when telemetry must never take down the service; keep False in development so misconfiguration is loud.

Environment-driven setup

Python
import os, opexia.trace

opexia.trace.init(
    org_id=os.environ["OPEXIA_ORG_ID"],
    workspace_id=os.environ["OPEXIA_WORKSPACE_ID"],
    project_id=os.environ.get("OPEXIA_PROJECT_ID", "default"),
    backend_url=os.environ.get("OPEXIA_INGEST_URL", "https://ingest.opexia.dev"),
    api_key=os.environ["OPEXIA_API_KEY"],
    transport=os.environ.get("OPEXIA_TRANSPORT", "direct"),
    fail_open=os.environ.get("ENV") == "production",
)
Shell
# .env
OPEXIA_ORG_ID=acme
OPEXIA_WORKSPACE_ID=3f2b1c8e-9a4d-4f11-b7c2-1e5a8d3f6b90
OPEXIA_PROJECT_ID=checkout-agent
OPEXIA_INGEST_URL=https://ingest.opexia.dev
OPEXIA_API_KEY=opx_live_xxxxxxxxxxxxxxxxxxxx
OPEXIA_TRANSPORT=direct

@observe

Decorate a function to emit one span per call. Works on both sync and async functions — the decorator detects which and wraps accordingly.

Python
from opexia.trace import observe

@observe(reasoning_role="retrieval", node_type="retrieval")
def fetch_order(order_id: str) -> dict:
    return db.get(order_id)

@observe(reasoning_role="synthesis", node_type="agent", name="compose_answer")
async def compose(ctx: dict) -> str:
    return await llm.complete(ctx)

Parameters

Keyword-only.

NameTypeRequiredDefaultDescription
reasoning_rolestr | None➖NoneWhat this step is for — see enums
node_typestr | None➖NoneWhat kind of component it is — see enums
end_userstr | None➖NoneEnd-user identifier for per-user attribution
namestr | None➖fn.__qualname__Span name override

Error handling

An exception inside the wrapped function is recorded on the span (record_exception), the span status is set to ERROR, and the exception is re-raised unchanged. Instrumentation never swallows your errors.


ReasoningTrace

A context manager for a logical unit of reasoning, giving you a node object to attach structured evidence to. This is what makes reconstruction possible — plain spans record what happened; these records capture why.

Python
from opexia.trace import ReasoningTrace

with ReasoningTrace("answer_customer", reasoning_role="synthesis",
                    node_type="agent", end_user="user_8412") as node:

    node.record_decision(
        selected="lookup_order",
        scores={"lookup_order": 0.88, "search_kb": 0.31},
        alternatives=[{"name": "search_kb", "why_not": "no order id in message"}],
        rules_fired=["order_id_present"],
    )

    node.record_sources(
        consulted=["https://docs.internal/returns", "https://docs.internal/shipping"],
        used=["https://docs.internal/returns"],
        dropped=["https://docs.internal/shipping"],
        scores={"https://docs.internal/returns": 0.93},
    )

    node.record_cost(
        model="claude-opus-4-6",
        input_tokens=2_140,
        output_tokens=310,
    )

Constructor

NameTypeRequiredDefaultDescription
namestr✅—Span name (positional)
domainstr | None➖NoneFree-form label. Not indexed — not queryable server-side.
reasoning_rolestr | None➖NoneSee enums
node_typestr | None➖NoneSee enums
end_userstr | None➖NoneEnd-user identifier

Node methods

record_decision(*, selected, scores=None, alternatives=None, rules_fired=None)

ParamTypeRequiredDescription
selectedstr✅The option chosen
scoresdict[str, float]➖Score per option
alternativeslist[dict]➖Options not taken, with rationale
rules_firedlist[str]➖Deterministic rules that applied

Feeds the decision-trace engine.

record_sources(*, consulted, used=None, dropped=None, scores=None)

ParamTypeRequiredDescription
consultedlist[str]✅Every source retrieved — string IDs or URLs, not objects
usedlist[str]➖The subset actually cited. Drives source_authority.
droppedlist[str]➖Retrieved and deliberately discarded
scoresdict[str, float]➖Relevance per source, 0–1

Feeds the sources matrix. The consulted vs used split is what surfaces "retrieved a good source and ignored it".

record_cost(*, model, input_tokens, output_tokens, ...)

ParamTypeRequiredDefaultDescription
modelstr✅—Model identifier
input_tokensint✅—Prompt tokens
output_tokensint✅—Completion tokens
reasoning_tokensint➖0Extended-thinking tokens
tool_call_countint➖0Tool invocations
injected_context_tokensint➖0Tokens placed into context
used_context_tokensint➖0Tokens the model actually drew on

USD and the pricing version are computed for you. The injected vs used context split powers context-optimization recommendations.

record_reliability_inputs(**kwargs)

Free-form signals for the reliability scorer.

record_plan(plan)

Records the agent's plan; feeds the decomposition engine.

record_end_user(end_user)

Sets opexia.end_user after construction.

subnode(...)

Creates a nested node under the current one.


Cost estimation

Python
from opexia.trace import estimate_cost_usd, pricing_version

usd = estimate_cost_usd(model="claude-opus-4-6", input_tokens=2140, output_tokens=310)
print(f"${usd:.4f} (pricing {pricing_version()})")

Use this for local pre-flight estimates. Server-side cost on stored traces is computed independently, so a stored figure never shifts because your SDK version changed.


Enumerations

Both are server-validated — any other value fails validation and the span is dead-lettered.

reasoning_role — what the step is for:

decomposer · research · analysis · critique · synthesis · arbiter · retrieval · guardrail · post_process

node_type — what kind of component it is:

decomposer · classifier · agent · guardrail · post_process · retrieval


Span attribute reference

The wire contract. Violations are silently dead-lettered at ingest — the span disappears rather than erroring, so get these right.

Required envelope

AttributeTypeDescription
opexia.schema_versionstringCurrently "1.0"
opexia.org_idstringOrganization identifier
opexia.workspace_idstringWorkspace UUID
opexia.project_idstringProject/service name
opexia.trace_idstringMust equal the OTel trace ID

Optional attributes

AttributeTypeDescription
opexia.user_idstringInternal user identifier
opexia.end_userstringEnd-user identifier for per-user attribution
opexia.reasoning_roleenumSee above
opexia.node_typeenumSee above
opexia.parent_reasoning_idstringParent reasoning node
opexia.decisionJSON string{rules_fired, scores, selected, alternatives}
opexia.sourcesJSON string{consulted, used, dropped, scores}
opexia.query_textplain stringThe user query / the prompt. ~16 KB cap.
opexia.outcome_textplain stringThe final answer. ~16 KB cap.
opexia.prompt_idstringPins prompt identity for Ship Check
opexia.prompt_labelstringDisplay name
opexia.prompt_versionstringClient-computed version
opexia.cost.usdfloatFlat scalar
opexia.cost.model_pricing_versionstringPricing table version

Standard OTel GenAI semantic conventions are also read: gen_ai.system, gen_ai.request.model, gen_ai.operation.name, gen_ai.usage.input_tokens, gen_ai.usage.output_tokens, gen_ai.response.finish_reason.

The four rules

  1. Compound fields are JSON strings. opexia.decision and opexia.sources must be json.dumps(...). A nested object becomes an OTLP kvlistValue and kills the whole span.
  2. consulted / used / dropped are string[] — URLs or stable IDs, never {id, title} objects.
  3. query_text / outcome_text are plain strings. Never json.dumps them.
  4. Only documented opexia.* keys exist (the schema is extra="forbid"). Any other opexia.* key kills the span. Note the keys are flat: opexia.prompt_id, not opexia.prompt.id.

Where to put text. Engines take the first non-empty query_text and the last non-empty outcome_text in a trace. Put the query on the earliest span and the outcome on the latest.

On an LLM span, query_text is the prompt — it is what Evigauge versions, evaluates, and measures a cacheable prefix from. When rendering messages into it, keep role boundaries:

Python
# ✅ Correct — role boundaries preserved
query_text = "\n\n".join(f"{m['role']}:\n{m['content']}" for m in messages)

# ❌ Wrong — no boundary, so every call hashes as its own "prompt" and the
#    Prompts page fills with thousands of one-call rows.
query_text = " ".join(m["content"] for m in messages)

Durability

Write-ahead log

Spans are written to wal_path (default .opexia-wal/spans.jsonl) before export. If the process crashes, the next init() replays the WAL and logs how many entries it recovered.

Replay happens before the span processor is constructed — required on Windows, where renaming a file with an open handle raises PermissionError (WinError 32).

Mount .opexia-wal/ on a writable volume in containers. On a fully read-only filesystem, point wal_path at a writable temp directory.

Sampling

Python
opexia.trace.init(..., sampler_rate=0.1)   # 10% of traces

Sampling is per trace, not per span — a sampled trace keeps all its spans, so reconstruction still works. Sampling a fraction of spans would produce broken traces.


Complete Python example

Python
"""Minimal end-to-end instrumented agent."""
import os
import opexia.trace
from opexia.trace import ReasoningTrace, observe

opexia.trace.init(
    org_id=os.environ["OPEXIA_ORG_ID"],
    workspace_id=os.environ["OPEXIA_WORKSPACE_ID"],
    project_id="checkout-agent",
    backend_url="https://ingest.opexia.dev",
    api_key=os.environ["OPEXIA_API_KEY"],
    transport="direct",
    fail_open=os.environ.get("ENV") == "production",
)


@observe(reasoning_role="retrieval", node_type="retrieval")
def retrieve(query: str) -> list[str]:
    return ["https://docs.internal/returns", "https://docs.internal/shipping"]


def handle(query: str, user_id: str) -> str:
    # One ReasoningTrace per logical request = one trace.
    with ReasoningTrace("handle_request", reasoning_role="synthesis",
                        node_type="agent", end_user=user_id) as node:

        # Put the query on the EARLIEST span of the trace.
        node._span.set_attribute("opexia.query_text", query)

        docs = retrieve(query)
        node.record_decision(
            selected="answer_from_kb",
            scores={"answer_from_kb": 0.91, "escalate": 0.12},
            rules_fired=["kb_hit"],
        )
        node.record_sources(consulted=docs, used=docs[:1], dropped=docs[1:])

        answer = f"Returns are accepted within 30 days. ({len(docs)} sources)"

        node.record_cost(model="claude-opus-4-6", input_tokens=2140, output_tokens=310)
        # Put the outcome on the LATEST span of the trace.
        node._span.set_attribute("opexia.outcome_text", answer)
        return answer


if __name__ == "__main__":
    print(handle("Can I return this?", user_id="user_8412"))

Verifying a span landed

Run this after wiring up instrumentation. It emits one span and confirms it arrived — the definitive check that the whole path works.

Python
"""verify_span.py — emit one span and confirm it landed."""
import os, time, uuid, httpx
import opexia.trace
from opexia.trace import ReasoningTrace

ORG   = os.environ["OPEXIA_ORG"]            # slug, for the read API
WS    = os.environ["OPEXIA_WORKSPACE"]      # slug, for the read API
WS_ID = os.environ["OPEXIA_WORKSPACE_ID"]   # UUID, for the SDK
KEY   = os.environ["OPEXIA_API_KEY"]

opexia.trace.init(
    org_id=ORG, workspace_id=WS_ID, project_id="verify",
    backend_url="https://ingest.opexia.dev", api_key=KEY, transport="direct",
)

marker = f"verify-{uuid.uuid4().hex[:8]}"
with ReasoningTrace("verify_span", reasoning_role="analysis", node_type="agent") as n:
    n._span.set_attribute("opexia.query_text", marker)
    n.record_cost(model="claude-opus-4-6", input_tokens=10, output_tokens=5)

# Force the batch out, then let ingestion settle.
from opentelemetry import trace as ot
ot.get_tracer_provider().force_flush()
print(f"emitted {marker}; waiting for ingestion…")
time.sleep(20)

# 1. Did anything dead-letter?
health = httpx.get(
    f"https://api.opexia.dev/v1/observ/orgs/{ORG}/workspaces/{WS}/health/ingestion",
    headers={"x-opexia-api-key": KEY}, timeout=30,
).json()
print("dead-lettered last hour:", health.get("dead_letter_count_last_hour"))

# 2. Is the trace queryable?
traces = httpx.get(
    f"https://api.opexia.dev/v1/observ/orgs/{ORG}/workspaces/{WS}/traces",
    headers={"x-opexia-api-key": KEY},
    params={"page_size": 20, "include_empty": True}, timeout=30,
).json()
# Payload is under "data", not "items".
print("recent traces:", len(traces["data"]))
assert traces["data"], "No traces — check key, workspace UUID, and ingest URL."
print("✅ span landed")

If nothing lands, work through this in order:

CheckHow
Key valid and pointed at the right workspace?GET https://ingest.opexia.dev/v1/auth/whoami
Spans arriving at all?spans_last_hour on ingestion health
Arriving but rejected?dead_letter_count_last_hour > 0 → an attribute breaks the four rules
Using a collector?Its otlphttp exporter must set encoding: json — protobuf returns 415
Process exiting immediately?Call force_flush() before exit, or spans die in the batch queue