dprovenance.dev / compare / dprovenancekit-vs-langsmith

DProvenanceKit vs LangSmith: catch the regression your evals miss

DProvenanceKit is regression testing for AI agent execution paths: record a golden decision path, diff the next run, fail CI when a required step disappears — even when the final answer still looks right. LangSmith is the hosted observability and evals platform you keep for dashboards, monitoring, and annotation. Use them together.

Most stacks already have output tests, LLM judges, and OpenTelemetry. Those can all stay green while the agent quietly drops verify. That gap is the whole product.

Same agent change: a required verify step is deleted. Final answer still looks fine.
Output test (string / snapshot)PASS
LLM-as-judge / eval suitePASS
OpenTelemetry / LangSmith monitoringPASS
DProvenanceKit execution-path gateFAIL — verify step disappeared
That is regression testing for AI agent execution paths — the front door. Cryptographic attestation is how you later prove the path that shipped. Run the 60-second CI demo →

# 01 · at a glance

Execution-path regression vs hosted observability

Every claim in the LangSmith column below was checked against langchain.com and docs.langchain.com on July 3, 2026. LangSmith's pricing and features move; treat their pricing page as the source of truth for current numbers.

 DProvenanceKitLangSmith
Tool shape Python SDK + CLI. pip install, import, done — runs entirely in-process, no service to stand up. Hosted observability + evals platform with SDK clients: tracing, dashboards, monitoring, insights, a managed backend.
Dependency footprint ✓ Zero third-party dependencies in the core — stdlib only (sqlite3, contextvars, hashlib, …). SDK client plus a running service — cloud or self-hosted — to send traces to.
Data locality ✓ Local SQLite files that live in your repo. Traces never leave your machine. Traces go to the hosted service by default; self-hosted and hybrid deployments are offered on the Enterprise plan (as of July 2026).
License ✓ Apache-2.0 — SDK, gate CLI, pytest plugin, and GitHub Action are all open source. SDK clients are MIT-licensed; the platform itself is a proprietary hosted service.
CI gate on agent execution path ✓ Built in, four ways: assert_no_regression, the golden_trace pytest fixture, the server-less gate CLI, and a GitHub Action (+ a GitLab CI template). The @pytest.mark.langsmith integration logs test results to the platform as datasets and experiments; its docs as of July 2026 don't describe a built-in structural golden-baseline gate.
Cryptographic attestation / audit proof ✓ Offline-verifiable decision-path attestation (Swift: Secure Enclave + proof packs; Python MVP via dprovenancekit[crypto] / DPK-BINARY-V1). Exportable evidence for compliance review — not a prettier chart. Does not claim regulator acceptance or model soundness. — Observability and evals platform; docs as of mid-2026 do not describe offline cryptographic attestation of an agent decision path as a product surface.
Framework coverage LangChain / LangGraph, OpenAI Agents SDK (officially listed in its docs), LlamaIndex, CrewAI, Google GenAI, FastAPI, MCP — or plain Python with zero deps. Framework-agnostic: LangChain / LangGraph natively, plus OpenAI SDK, Anthropic SDK, Vercel AI SDK, LlamaIndex, and OpenTelemetry ingestion (as of July 2026).
Pricing model ✓ SDK, CLI, pytest plugin, Action, and the in-browser trace Explorer: all free and open source. Free Developer tier (1 seat, up to 5k base traces/mo), Plus at $39/seat/mo, usage-based trace pricing beyond included volume, custom-priced Enterprise (as of July 2026).
Languages Python 3.9+, plus a Swift SDK kept behaviorally equivalent through Trace Specification v1. Python and TypeScript SDKs (MIT); langchain.com also lists Go and Java SDKs as of July 2026.

Sources: langchain.com/langsmith, langchain.com/pricing, docs.langchain.com/langsmith, github.com/langchain-ai/langsmith-sdk — all retrieved July 3, 2026.

# 02 · the core difference

Regression testing first; attestation when buyers ask for proof

The category at the front door is AI agent execution-path regression testing. Golden path → record → change prompt/model/tools → record candidate → structural + semantic diff → CI fails when verify vanished. Attestation and offline proof packs are the deeper layer for regulated and enterprise buyers — not the first sentence on the box.

  • The killer demo is the differentiator; the Action is the distribution. Output tests, evals, and OTel can all pass while a required step disappears. Fork the synthetic regression demo and watch CI fail on a dropped verify. Layer Swift Secure Enclave / Python [crypto] attestation when a buyer asks for offline proof.
  • Golden baselines are files. A known-good run is saved as a SQLite trace (e.g. tests/goldens/research-agent.sqlite) and committed like any other snapshot. pytest --dprov-update-golden re-records it deliberately; a plain pytest gates against it.
  • Run fingerprints make drift cheap to detect. Every run gets a structural identity of its execution path — a reordered tool, a skipped retrieval step, a dropped verify step all change the fingerprint.
  • Diffs are severity-graded, not binary. Regressions grade none / low / medium / high; a removed critical step (think verify.claims or safety_check) grades high. The gate is strict by default and loosens explicitly via max_regression_level or allow_divergent_steps, and you can plug in a custom evaluator.
  • Gating is built in at every layer — assert_no_regression in any test, the pytest fixture below, the server-less dprovenancekit gate CLI (local SQLite, no backend, no account), and the GitHub Action with a GitLab CI template in the repo.
tests/test_agent.pypython
def test_research_agent(golden_trace):
    with golden_trace("research-agent"):
        run_my_agent()

# pytest --dprov-update-golden   -> records tests/goldens/research-agent.sqlite
# pytest                         -> gates against it; fails on structural drift

New to the golden-baseline pattern? The concepts are in the regression testing guide; the pipeline wiring is in the CI gate guide.

# 03 · five minutes

The whole loop: record, save, explain, diff

This is the part where "free, local-first" stops being an abstraction. One import, a context manager, and two method calls — no account, no API key, no service. Straight from the README, unedited:

five_minute_wow.pypython
from dprovenancekit import trace

# 1. Record an execution
with trace("Agent Workflow"):
    with trace("Retrieve Documents"):
        # your retrieval code here
        pass
    with trace("Verify Claims"):
        # your verification code here
        pass

# 2. Save the trace
trace.save("golden_run.sqlite")

# 3. Print a structural explanation
trace.explain()
# --- Execution Trace (b4f8d2…) ---
# ▶ Started Agent Workflow
#   ▶ Started Retrieve Documents
#   ✔ Finished Retrieve Documents
#   ▶ Started Verify Claims
#   ✔ Finished Verify Claims
# ✔ Finished Agent Workflow

# 4. Later, the workflow regresses: a code change drops the verification step
with trace("Agent Workflow"):
    with trace("Retrieve Documents"):
        pass
    # "Verify Claims" never runs

# 5. Diff the current run against the saved golden baseline
trace.diff("golden_run.sqlite")
# --- Trace Diff (Golden vs Current) ---
# ❌ Missing step: Verify Claims

That last line is the entire pitch. The same diff engine powers the gate CLI whose real failing output — severity, fingerprints, per-step changes, exit code 1 — is captured on the homepage demo.

# 04 · not mutually exclusive

Gate LangChain agents while keeping LangSmith

Here is the part vendor comparison pages usually skip: you don't have to choose. DProvenanceKit ships a LangChain adapter — DProvenanceTracer and DProvenanceCallbackHandler — and the handler is a normal LangChain callback, so it runs alongside LangSmith's own tracing in the same invocation. Keep LangSmith for the dashboards it is good at, and add a structural regression gate to the same LangChain / LangGraph agents in CI:

gate_langchain_agent.pypython
# pip install "dprovenancekit[langchain]"
from dprovenancekit import SQLiteTraceStore
from dprovenancekit.integrations.langchain import DProvenanceTracer, LangChainTraceEvent

store = SQLiteTraceStore(LangChainTraceEvent, "traces.sqlite")
tracer = DProvenanceTracer(store)

with tracer.trace(context_id="customer-42") as cb:
    result = chain.invoke(question, config={"callbacks": [cb]})

# afterwards: query / diff / fingerprint the recorded run via `store`.

Same story for the OpenAI Agents SDK, where DProvenanceKit is officially listed in the SDK's docs as an external tracing processor — one register() call, and it coexists with the SDK's built-in exporter.

# 05 · validation

Small tool, serious test surface

A regression gate you can't trust is worse than no gate. The detection engine is validated three ways:

  • 380+ tests cover the SDK, the diff engine, the pytest plugin, the CLI, and the adapters.
  • Precision 1.000 / Recall 1.000 / F1 1.000 across 8 standard + 5 adversarial regression-detection scenarios, on a benchmark corpus that ships with the library — run dprovenancekit evaluate yourself; no vendor benchmark server involved.
  • Trace Specification v1 — a conformance suite with frozen golden vectors keeps the Python SDK behaviorally equivalent to the original Swift implementation, case-for-case.

# 06 · the other column

When LangSmith is the better choice

This is a genuine section, not a formality. As of July 2026, LangSmith is the better tool when:

  • Your team lives in a shared dashboard. LangSmith's monitoring dashboards, alerting, trace search, and cross-run analytics are a real product; DProvenanceKit has no hosted dashboard — it's a local-first library.
  • You want hosted evals with humans in the loop. Annotation queues, inline annotation, online LLM-as-judge evaluations, and datasets/experiments are LangSmith platform features with no DProvenanceKit equivalent.
  • Your stack isn't Python or Swift. LangSmith ships Python and TypeScript SDKs (with Go and Java also listed on langchain.com); DProvenanceKit's cross-language story is exactly Python + Swift via Trace Spec v1, no further.
  • You want a managed service, not a library. If nobody on the team should own trace storage or report plumbing, a hosted platform with a free Developer tier is the pragmatic call.

If several of those describe you, use LangSmith — and consider running the DProvenanceKit gate next to it, since the two coexist in the same agent (see # 04).

# 07 · get started

Reproduce the FAIL in sixty seconds

The honest test of any complement to LangSmith claim is what happens in the first ten minutes. Record one golden run, break the agent on purpose, and watch the diff fail. No account, nothing leaves your machine.

shell
$ pip install dprovenancekit

View on GitHub → PyPI · GitHub Action

Then: the regression testing guide for the concepts, and the CI gate guide to wire it into CI.

This comparison was written by the DProvenanceKit maintainer; LangSmith specifics were verified against langchain.com and docs.langchain.com on July 3, 2026. Spotted something stale or unfair? Corrections are welcome via GitHub issue.