Motivation / problem
The v4.7 Tabular Review engine turned a set of documents into a spreadsheet: rows are documents, columns are questions, and each cell is a RAG-grounded extraction with a flag (green/yellow/red) and cited chunks. It answers “what does each document say about X?” well. But every column was the same kind of question — a single-shot LLM extraction — so two things were out of reach:- Governance questions that the graph already knows the answer to. “Is this doc canonical? Orphaned in the graph? Superseded? What’s its evidence tier?” are deterministic facts in the canonical layer, not LLM judgments. Asking an LLM to guess them is slower, costs money, and can hallucinate.
- Anti-hallucination on the extraction itself. A green flag means the model was confident — not that the value is actually supported by the cited evidence.
agent dimension orthogonal to its format, so a
single report can mix RAG extraction, deterministic graph governance, and a verified
anti-hallucination pass — and ship as a one-click, per-cell-cited, exportable audit
matrix.
v8.40 (W6, ADR 0034) adds a fourth kind, vision,
for a third class of question the original three couldn’t answer: questions the
text doesn’t have the answer to. “What colour is this garment? What material?
What pattern?” — nothing a fashion-ecommerce catalog’s OCR’d text says answers
this; the answer is only visible in the product photo.
Theory & background
The unit of the engine is still a cell = (document row × column). What changes is that a column declares how its cell is produced, via theagent kind
(App\Support\TabularReview\AgentKind):
extract— today’s behaviour, unchanged and the default. One RAG single-shot per (doc, column): retrieve, ask the LLM, flag + cite. Every pre-v8.19 review is byte-identical (the field is absent →extract).graph— a deterministic, LLM-free governance metric resolved from the canonical graph + document columns. No model call, no cost, no hallucination — the value is a fact (is_canonical = true,incoming_edges = 4, …). Thegraphagent wins over thejson_pathshortcut when both are set.verify— a bounded anti-hallucination second pass. After anextractproduces a value, a verify call re-checks it against the document’s cited evidence and can only downgrade the flag (green → yellow / else → red) when the value isn’t supported. It is never worse than extract: a verify-call failure keeps the original cell (R14).vision(v8.40, ADR 0034) — one vision-LLM call per document, unbatched (likegraph, one resolver call per column per document — never part of theextract/verifybatched call, because the visual modality means the call shape is fundamentally per-document, not per-report). Nometrickey; the column’s existingprompt/format/enum_valuesdrive the extraction instruction exactly as they do forextract. Likegraph,visionwins over thejson_pathshortcut when both are set.
graph) → grounded
(extract) → grounded-and-checked (verify) → seeing (vision) — expressed as a
per-column choice rather than a whole-report mode.
The governance metrics come from one place — App\Services\TabularReview\GovernanceColumnResolver
— which reads the real taxonomies (EvidenceTier, CanonicalStatus enums) and
the canonical graph (kb_edges), tenant-scoped (R30). It exposes 10 metrics:
Design
The router lives inApp\Services\TabularReview\TabularReviewExtractor, which now
takes both GovernanceColumnResolver and VisionColumnResolver as constructor
dependencies. Per column, it picks a path by agent (graph and vision both win
over json_path, both LLM-free-or-per-document; otherwise extract, optionally
followed by verify):
The verify pass is bounded and order-stable: it ksorts the columns it checks and
only ever downgrades, so a report’s verified run is deterministic and can never invent
a better flag than the extraction earned.
Vision columns — image-source precedence
VisionColumnResolver never asks the admin to point at an image explicitly — it
derives the image(s) from the document itself, in a fixed order (ADR 0034 §2):
- OCR-extracted figures (
OcrFigureStore, the OCR guide) — read viaOcrService::status($doc), the same tenant-scoped call the Digitization Review UI already uses, forfigures_dir+figurescount. Only themistralanddoclingOCR drivers extract figures;vision-llmOCR extracts none (a named, pre-existing limitation this feature doesn’t change). - The document’s own source file, when it is itself an image — a standalone product photo ingested as one KB document (the primary shape for most fashion-ecommerce catalogs: one photo per SKU, not embedded in a PDF).
- Neither → a definite red cell, no provider call, no invented answer (R14).
StorageNamespace — the same tenant/project
storage-namespace resolution every KB read uses — so a vision column can never cross
a tenant’s storage boundary. Images are capped at
KB_TABULAR_VISION_MAX_IMAGES (default 4) per cell — one catalog page with many
figures never turns one cell’s generation into an unbounded provider call
(SEC-LLM-001 gate 7).
The provider call itself reuses the SdkAnonymousAgent + Base64Image pattern the
OCR guide’s vision-llm driver already established — metered automatically by the
laravel/ai SDK lifecycle hook, no double-counting, no new infrastructure.
The flagship “Canonical KB Governance Audit” preset (#16 in the seeded library)
is the headline application: rows = canonical docs, columns = 8 graph governance
auditors + 1 verify contradiction check. Running it turns the whole KB into a
per-cell-cited, exportable governance matrix — “which canonical docs are orphaned,
stale, superseded, low-evidence-tier, or self-contradictory?” answered in one grid.
The FE (built on the existing accessible DOM matrix) adds three surfaces: an agentic
column editor (the governance metric picker appears only for graph columns; submit
is gated so a graph-without-metric can’t 422), a per-cell evidence side-panel
(summary + flag + reasoning + cited chunks), and a one-click template gallery of
the built-in system reports.
Data model / contract
The engine reuses the v4.7 tables — no new tables for the agentic upgrade. The agentic dimension is additive on the column config:
The ready-made library reuses
workflows (type = tabular, is_system = true) —
BuiltInWorkflowSeeder now mints 16 system templates (the +1 is the governance
preset), each with a sensible columns_config. Both agent and metric are validated
on Store/UpdateTabularReviewRequest (metric is required_if the agent is graph
and must be one of the 10 GovernanceColumnResolver::METRICS).
The capability is reachable on all three R44 surfaces over one shared core:
- PHP —
TabularReviewExtractor+GovernanceColumnResolver+VisionColumnResolver+ theBuiltInWorkflowSeederlibrary. - HTTP — the existing
api/admin/tabular-reviews/*(list / show / create / generate), now accepting agentic columns.visionneeded zero new validation beyond the enum case —Rule::in(AgentKind::values())already covered it. - MCP —
App\Mcp\Tools\KbRunReportToolon theenterprise-kbserver (read a saved report’s matrix, tenant-scoped R30, OFF-path safe R43), bounded byCOUNT(DISTINCT) + LIMITso an agent can’t pull an unbounded matrix. Avisioncolumn’sagentfield passes through the same projection unchanged — confirmed by a dedicated test, not assumed.
Security & flags (R32 / R30 / R43)
- RBAC. Every tabular-review route is behind
can:viewTabularReviews/can:manageTabularReviews, R32-matrix-locked. The agentic upgrade adds no new route group — it extends the existing controllers’ request contracts. - Tenant isolation (R30).
GovernanceColumnResolverscopes everykb_edges/ document query to the active tenant;VisionColumnResolverresolves images exclusively throughStorageNamespace::diskOf()/recordedPrefix()— the same namespace resolution every other KB storage read already goes through — so a vision cell can never read another tenant’s figures or source file;KbRunReportToolfilters the report + cells by tenant. A client cannot read another tenant’s report or widen scope through a filter. - OFF-path safe (R43).
KbRunReportToolreturns a well-formed empty payload when the report/cells are absent rather than throwing;verifyfailures andgraphunknowns degrade to a grey/red cell (R14) instead of erroring the whole generate.KB_TABULAR_VISION_ENABLEDships default OFF: avisioncolumn with the flag off produces a definite redfailedcell naming the flag — every other column in the same review still generates normally, never a 500, never a silent fallback toextractsemantics. - Bounded provider work (SEC-LLM-001 gate 7).
KB_TABULAR_VISION_MAX_IMAGES(default 4) caps how many images one cell’s provider call can carry — a catalog page with many figures cannot turn one cell into an unbounded multi-image call. The raw provider exception isLog::warning’d, never persisted into the cell’sreasoning(SEC-ERRLEAK-equivalent).
Decision rationale (ADR-style)
- A per-column
agentdimension, not a report-level mode. Governance audits mix deterministic facts and grounded judgments in the same grid; a whole-report mode would force a false choice. Makingagenta column property lets one report carrygraphfacts next toextract/verifyanswers. See architecture decisions. - Deterministic graph metrics over LLM guesses. Canonical status, edge degree,
orphan-ness and supersession are facts in the graph. Resolving them with
GovernanceColumnResolver(no model call) is faster, free, and cannot hallucinate — and it reuses the realEvidenceTier/CanonicalStatusenums so the taxonomy never drifts (R9). verifycan only downgrade. An anti-hallucination pass that could raise a flag would itself be a hallucination risk. Constraining it to downgrade-only (and keeping the original cell on failure) makes the verified run monotonic and never worse than the extraction (R14).- Reuse
workflowsfor the library. The ready-made “precotte” templates are just system-owned tabular workflows — no new model, no new admin surface, and they show up in the existing workflow gallery for free. - Glide canvas grid deferred. The accessible DOM matrix (per-cell testids + ARIA, R11/R15) ships now; the canvas grid + SSE progressive-paint are documented v8.19.x follow-ups because canvas cells can’t carry per-cell testids/ARIA — testability and a11y come first.
- Image source is derived, not admin-configured. A
visioncolumn doesn’t ask the admin to point at a specific image — it derives the image(s) from the document itself (OCR figures, then a standalone image document). A per-column “pick an image” UI would be a materially bigger surface (an image picker, upload flow, multi-image gallery model) for a case the derivation already covers for both realistic shapes of a fashion-ecommerce catalog. - No manual cell-value override. “Human review in the grid” is satisfied by the
existing review UX — flag colour, evidence panel, regenerate cell — the same
loop every other agent kind already has. Editing a persisted cell’s value directly,
without re-running extraction, would be a materially larger capability spanning
every
AgentKind, not a vision-specific one, and is explicitly deferred (ADR 0034 §6).
Worked example
Create a governance report from the flagship preset, then read its matrix over MCP. Agraph column is computed with no LLM call:
verify column downgrades a too-confident extraction:
enterprise-kb
server calls KbRunReportTool and gets the tenant’s report matrix, bounded by
COUNT(DISTINCT) + LIMIT:
/app/admin/tabular-reviews, “From template” opens the gallery,
picks “Canonical KB Governance Audit”, and pre-fills the create dialog; clicking a
populated cell opens the evidence side-panel with the summary, flag, reasoning and
the cited KB chunks.
A vision column on a fashion catalog, KB_TABULAR_VISION_ENABLED=true:
TabularReviewsList.tsx), since the value
being cited is a filename the model was shown, not a retrieved chunk.
Gotchas & operations
graphneeds ametric. Agraphcolumn without a governance metric is a misconfiguration — the FE gates submit on it and the BE validator rejects it (required_if), so a bad column can’t reach generation and 422.graphwins overjson_path. If a column sets both, the deterministic graph resolver runs;json_pathis the LLM-free shortcut forextract-family columns.verifyadds a bounded second call. It is the one agent kind that costs extra per cell (one verify call after the extract). Use it for the columns where anti-hallucination matters, not blanket-on every column.- Existing reviews are unchanged. A pre-v8.19 review has no
agenton its columns → every column isextract→ byte-identical output. The agentic upgrade is purely additive (R27). - Re-run after canonical changes.
graphcells reflect the canonical graph at generate time; promote/supersede/delete a doc and re-generate to refresh the governance matrix. visionneeds OCR figures or a standalone image document. A tenant using only thevision-llmOCR driver (which extracts no figures by design) and no standalone image ingestion gets an honest “no visual evidence” red cell for every vision column — never a fabricated answer. Switch tomistralordoclingOCR (KB_OCR_DRIVER) to extract figures, or ingest product photos as standalone documents.visionis one call per cell, not batched. A 100-row × 3-vision-column review generating from scratch issues up to 300 bounded provider calls, same shape asgraph’s one-call-per-cell, notextract’s one-call-per-document batching. Real cost, bounded byKB_TABULAR_VISION_MAX_IMAGESand metered automatically.
Canonical & promotion
The canonical graph (kb_edges, status, evidence tiers) the graph columns read.
Grounding & evidence tiers
The EvidenceTier taxonomy a governance audit ranks each doc by.
Multi-tenant isolation
The TenantContext every report, cell and graph query is scoped to (R30).
MCP server
The KbRunReportTool reader on the enterprise-kb server.