Skip to content

Compatibility & Benchmark Evidence

This page records public compatibility-smoke evidence and deterministic benchmark summaries for llmwiki-serve. It is not quality certification, runtime certification, upstream endorsement, or a public answer-quality claim.

Package registry status is tracked in Release Status & Compatibility. This page does not restate the current package baseline or claim package quality for that baseline. The smoke report summary below is compatibility-only; the deterministic benchmark rows remain strict quality-gate records and do not create a public quality claim.

0.2.9 Lexical-Default Regression Gate

Before llmwiki-serve==0.2.9 was published, the default lexical retrieval path was compared against released version 0.2.6 with no vector, hybrid, or agent-guided mode enabled. The adoption signal is practical: users can upgrade to 0.2.9 for the new integration surface while preserving default lexical quality in this gate.

DatasetQueriesVersionnDCG@10Recall@100MAP@100qrel-positive Precision@10Negative FPLatency p50 / p95 ms
SciFact full3000.2.60.60231093750.82744444440.56385999410.0800000000N/A552.942 / 848.620
SciFact full3000.2.90.60231093750.82744444440.56385999410.0800000000N/A481.893 / 730.956
NoMIRACL-ko judged pool4260.2.60.55197532280.97887323940.49256252090.17136150230.9530516432116.242 / 368.255
NoMIRACL-ko judged pool4260.2.90.55197532280.97887323940.49256252090.17136150230.9530516432100.144 / 386.182

Quality metrics matched exactly for both datasets in the retained summaries. SciFact latency was lower at both p50 and p95. Korean judged-pool p50 latency was lower, while p95 increased by 17.927 ms (+4.9%), so that remains a monitoring signal rather than a release blocker.

Limitations: this was a Windows local run with two repetitions; NoMIRACL-ko is a judged-pool fixture, not a full-corpus Korean benchmark; raw per-query rankings and latency vectors were not retained; and no OS cold-cache, superiority, model-answer quality, or vendor-runtime claim is made. Vector, hybrid, and agent-guided candidate-only arms are separate engineering evidence and are not used here as a released-vs-released superiority claim.

The compatibility-smoke input summarized here is a Windows upstream compatibility-smoke report with schema llmwiki-serve-upstream-candidate-smoke-v1, mode smoke, and evidence track compatibility-smoke. A stable tracked public report artifact is pending; this page intentionally does not link local or untracked report paths.

The deterministic benchmark summary also records Windows and Ubuntu/DGX parity for OpenWiki and Pratiyush. The Ubuntu/DGX run used Ubuntu 24.04 on aarch64 with the NVIDIA GB10 class hardware bucket / DGX Spark. OpenWiki and Pratiyush semantic metrics and payload-token metrics match the Windows deterministic runs.

Summary

SignalValue
Compatibility cases12
Source files projected1,397
Pages projected501
Approved pages projected501
Graph nodes projected2,092
Graph edges projected2,887
Mutation statusAll cases recorded unchanged checkout status and unchanged source hash.
License evidence7 cases reported an SPDX-style license value; 5 cases require license review because the report did not find an explicit repository content license at the pinned commit.
Deterministic parityOpenWiki and Pratiyush semantic and payload-token metrics match across Windows and Ubuntu/DGX deterministic runs.

All cases met their per-case smoke minimums for page counts and graph counts in the report. The smoke checks projection compatibility for pinned public static sources. It does not judge source content quality, retrieval quality, answer correctness, model behavior, or vendor-runtime conformance. No quality metrics are guessed on this page.

Environment Coverage

EnvironmentEvidence status
WindowsVerified by the current compatibility-smoke report and the finalized retrieval metric rows below.
UbuntuDeterministic parity verified on Ubuntu 24.04, aarch64; OpenWiki and Pratiyush semantic and payload-token metrics match Windows deterministic runs.
NVIDIA DGXDeterministic parity verified in the NVIDIA GB10 class hardware bucket / DGX Spark. This is not Qwen Agent tier validation.
macOSNot tested.

Compatibility Cases

CaseProductPinned commitAdapterSource files / pages / approved pagesGraph nodes / edgesLicense evidenceMutation status
atomic-compiler-basicatomicstrata/llm-wiki-compiler69701f609ae166e9da194c2d340699eb43abf77ellmwiki-markdown8 / 8 / 865 / 96MITClean
samuraigpt-agentSamurAIGPT/llm-wiki-agent11f66f1166994b35de2d7d3d0b246cb28847bbf2llmwiki-markdown30 / 3 / 311 / 9MITClean
pratiyush-llm-wikiPratiyush/llm-wikib1088890ee0743810a92577aecad946c6b3eb2d2llmwiki-markdown563 / 22 / 2286 / 108MITClean
logseq-exporter-test-graphlogseq/logseqa9a67f61ab29972d2e2b6c7a5864e6e3306c0d9alogseq73 / 54 / 54105 / 73AGPL-3.0Clean
foam-templatefoambubble/foam-template84fa1844270d214520aca32c01d4e27c6728d12efoam116 / 79 / 79785 / 940needs-review: no explicit repository content license found at pinned commitClean
dendron-test-workspacedendronhq/dendron4420715a421756518863c47005c8c49a38e37621dendron202 / 154 / 154403 / 440Apache-2.0Clean
karpathy-llm-wiki-vaultjason-effi-lab/karpathy-llm-wiki-vault18f4e71518af7d0c51a2fc65f5e3ec3043668e54llmwiki-markdown19 / 19 / 19163 / 315needs-review: no explicit repository content license found at pinned commitClean
langchain-openwiki-self-docslangchain-ai/openwiki9c253af17f264ac2589ab6781e79e9bb5b5d1238llmwiki-markdown14 / 13 / 13116 / 125MITClean
microsoft-llmwiki-fixturesmicrosoft/llmwiki74a8a5bf0011b1092f135e5cbc51bbb44c1e07e7generic-markdown5 / 5 / 513 / 8MITClean
luotwo-llm-wikiluotwo/llm-wiki9ab20ee0e9db3ca0bc7998b1b4a97ba7c821279fllmwiki-markdown15 / 11 / 1151 / 95needs-review: no explicit repository content license found at pinned commitClean
nishio-llm-wiki-about-delitenishio/llm-wiki-about-delite4181dd42ff78d72a5e5a05512a59dc37d7ef97a2quartz318 / 129 / 129261 / 648needs-review: no explicit repository content license found at pinned commit; Quartz/tooling license is not treated as sampled content licenseClean
iblinkq-llm-wiki-obsidian-blinkiBlinkQ/llm-wiki-obsidian-blinka9e8399cc29dbcce75fb47f61f1f2034a9dfc199obsidian34 / 4 / 433 / 30needs-review: no explicit repository content license found at pinned commitClean

Clean means the smoke report recorded both checkout_status_unchanged: true and source_hash_unchanged: true for that case.

Sidecar Boundary

Native hot.md, index.md, overview.md, and root quickstart.md files are untouched. The managed-context sidecar is a generic-only external opaque sidecar outside the served root. It is not a source page or graph file.

The sidecar boundary keeps managed-context metadata separate from native source content. Clients should continue to treat source pages and graph files as the served projection inputs, and should not infer local files or private storage layout from opaque handles.

Retrieval Quality

Deterministic retrieval metrics are recorded for the verified OpenWiki and Pratiyush summaries. Ubuntu/DGX parity is verified on Ubuntu 24.04, aarch64, NVIDIA GB10 class hardware bucket / DGX Spark; semantic metrics and payload-token metrics match the Windows deterministic runs, so separate metric tables are not duplicated here. public_quality_claim=false for every listed row, and no quality pass is claimed.

Tokenizer accounting used Qwen/Qwen3.6-35B-A3B, revision 53c43178507d69762986fbfa314f6e8d4d859409, with Qwen2Tokenizer. The reports record local Qwen tokenizer-load verification and no byte/mock token-count proxy.

qrel-positive metrics below are judged-positive retrieval metrics over returned corpus or page IDs. They are not citation metadata validation; true source-reference and citation-contract checks are described separately where they are explicitly named.

CaseVariantCorpus recordsQueriesQrelsGatepublic_quality_claim
OpenWikinative1357164failfalse
OpenWikigeneric-shadow1355141failfalse
Pratiyushnative2260172failfalse
Pratiyushgeneric-shadow2255144failfalse

Native cold evidence rows use service_context runs. Native managed-on evidence is an exact no-op versus managed off for these reports: ranking, qrel-positive precision/recall metrics, negative FPR, and payload tokens are unchanged.

CaseManagedRecall@5MRRnDCG@10qrel-positive Precision/RecallNegative FPRPayload p50/p95 tokensGate
OpenWikioff0.91690.90540.84240.4877 / 0.87421.00008400 / 8410fail
OpenWikion0.91690.90540.84240.4877 / 0.87421.00008400 / 8410fail
Pratiyushoff0.85270.88150.80060.3986 / 0.74361.00007390 / 7421fail
Pratiyushon0.85270.88150.80060.3986 / 0.74361.00007390 / 7421fail

Managed Context On/Off

Generic-shadow cold evidence rows use service_context runs. Managed-on evidence is unchanged versus managed off; only payload tokens changed.

CaseManagedRecall@5MRRnDCG@10qrel-positive Precision/RecallNegative FPRPayload p50/p95 tokensGate
OpenWikioff0.99600.96670.92000.4909 / 0.99261.00007386 / 7397fail
OpenWikion0.99600.96670.92000.4909 / 0.99261.00007319 / 7330fail
Pratiyushoff0.85700.85330.80370.3647 / 0.74051.00006963 / 6991fail
Pratiyushon0.85700.85330.80370.3647 / 0.74051.00006892 / 6928fail
CaseEvidence deltaSamples
OpenWikievidence unchanged; token delta mean -67, CI95 [-67, -67]1000
Pratiyushevidence unchanged; token delta mean -71.3, CI95 [-77.7, -65.6]1000

Generic-shadow cold orientation rows use service_context_orientation runs. Managed context improved orientation retrieval, but negative FPR stayed at 1.0000 and all quality gates still failed.

CaseManagedRecall@5MRRnDCG@10qrel-positive Precisionqrel-positive RecallNegative FPRPayload p50/p95 tokens
OpenWikioff0.15270.16000.09550.14550.17651.00007386 / 7397
OpenWikion0.51530.72000.51400.60910.49261.00007319 / 7330
Pratiyushoff0.00000.00000.00000.00000.00001.00006963 / 6991
Pratiyushon0.55400.77000.53000.55140.45041.00006892 / 6928

OpenWiki orientation point deltas were recall +0.3627, MRR +0.5600, and nDCG +0.4185. The verified OpenWiki report does not emit paired off/on orientation bootstrap CI, so none is claimed.

Pratiyush orientation deltas with CI95 and 1000 samples were recall +0.5540 CI [0.4607, 0.6457], MRR +0.7700 CI [0.6700, 0.8600], nDCG +0.5300 CI [0.4519, 0.6024], and payload tokens mean -188.3 CI [-427.9, -69.1].

Cold latency changed across runs and is noisy. This page makes no latency improvement or public latency claim. Pratiyush evidence cold latency improved in CI, but OpenWiki evidence and orientation cold latency worsened.

No public retrieval-quality claim is made. The current blockers are: negative FPR is 1.0000 for all listed cold evidence and orientation rows, above the <=0.05 threshold; qrel-positive precision is below 0.95 in all listed evidence and orientation rows; OpenWiki native evidence nDCG@10 is 0.8424, below 0.85; Pratiyush native and generic-shadow evidence miss recall and/or nDCG/qrel-positive recall thresholds as listed by report gates; selected search-read token p95 gates still fail in the verified reports; Qwen agent-tier validation is pending; and macOS remains untested.

Qwen Agent Tier

Qwen Agent tier metrics remain pending. The deterministic retrieval summaries verify tokenizer accounting for Qwen payload measurements on Windows and Ubuntu/DGX, but they do not verify Qwen Agent tool use, tool-call failure rate, source-report integrity, citation integrity, or model-answer quality.

No public Qwen Agent quality or tier claim is made.

Non-Claims

  • Not quality certification.
  • Not MCP, A2A, Qwen Agent, or vendor-runtime certification.
  • Not a model-answer quality benchmark.
  • Not an endorsement or affiliation claim for any upstream producer.
  • Not proof that private wiki content is safe to expose without operator review, network controls, authentication, and logging policy.

Public-preview documentation for Knowledge Bridge Labs wiki Knowledge Source components.