Skip to main content

Motivation / problem

The v4.7 Tabular Review engine turned a set of documents into a spreadsheet: rows are documents, columns are questions, and each cell is a RAG-grounded extraction with a flag (green/yellow/red) and cited chunks. It answers “what does each document say about X?” well. But every column was the same kind of question — a single-shot LLM extraction — so two things were out of reach:
  1. Governance questions that the graph already knows the answer to. “Is this doc canonical? Orphaned in the graph? Superseded? What’s its evidence tier?” are deterministic facts in the canonical layer, not LLM judgments. Asking an LLM to guess them is slower, costs money, and can hallucinate.
  2. Anti-hallucination on the extraction itself. A green flag means the model was confident — not that the value is actually supported by the cited evidence.
Agentic Knowledge Reports (v8.19) promotes the engine to first-class agentic columns: a column now has an agent dimension orthogonal to its format, so a single report can mix RAG extraction, deterministic graph governance, and a verified anti-hallucination pass — and ship as a one-click, per-cell-cited, exportable audit matrix. v8.40 (W6, ADR 0034) adds a fourth kind, vision, for a third class of question the original three couldn’t answer: questions the text doesn’t have the answer to. “What colour is this garment? What material? What pattern?” — nothing a fashion-ecommerce catalog’s OCR’d text says answers this; the answer is only visible in the product photo.

Theory & background

The unit of the engine is still a cell = (document row × column). What changes is that a column declares how its cell is produced, via the agent kind (App\Support\TabularReview\AgentKind):
  • extract — today’s behaviour, unchanged and the default. One RAG single-shot per (doc, column): retrieve, ask the LLM, flag + cite. Every pre-v8.19 review is byte-identical (the field is absent → extract).
  • graph — a deterministic, LLM-free governance metric resolved from the canonical graph + document columns. No model call, no cost, no hallucination — the value is a fact (is_canonical = true, incoming_edges = 4, …). The graph agent wins over the json_path shortcut when both are set.
  • verify — a bounded anti-hallucination second pass. After an extract produces a value, a verify call re-checks it against the document’s cited evidence and can only downgrade the flag (green → yellow / else → red) when the value isn’t supported. It is never worse than extract: a verify-call failure keeps the original cell (R14).
  • vision (v8.40, ADR 0034) — one vision-LLM call per document, unbatched (like graph, one resolver call per column per document — never part of the extract/verify batched call, because the visual modality means the call shape is fundamentally per-document, not per-report). No metric key; the column’s existing prompt/format/enum_values drive the extraction instruction exactly as they do for extract. Like graph, vision wins over the json_path shortcut when both are set.
This is the agentic ladder — cheap-and-certain (graph) → grounded (extract) → grounded-and-checked (verify) → seeing (vision) — expressed as a per-column choice rather than a whole-report mode. The governance metrics come from one place — App\Services\TabularReview\GovernanceColumnResolver — which reads the real taxonomies (EvidenceTier, CanonicalStatus enums) and the canonical graph (kb_edges), tenant-scoped (R30). It exposes 10 metrics:

Design

The router lives in App\Services\TabularReview\TabularReviewExtractor, which now takes both GovernanceColumnResolver and VisionColumnResolver as constructor dependencies. Per column, it picks a path by agent (graph and vision both win over json_path, both LLM-free-or-per-document; otherwise extract, optionally followed by verify): The verify pass is bounded and order-stable: it ksorts the columns it checks and only ever downgrades, so a report’s verified run is deterministic and can never invent a better flag than the extraction earned.

Vision columns — image-source precedence

VisionColumnResolver never asks the admin to point at an image explicitly — it derives the image(s) from the document itself, in a fixed order (ADR 0034 §2):
  1. OCR-extracted figures (OcrFigureStore, the OCR guide) — read via OcrService::status($doc), the same tenant-scoped call the Digitization Review UI already uses, for figures_dir + figures count. Only the mistral and docling OCR drivers extract figures; vision-llm OCR extracts none (a named, pre-existing limitation this feature doesn’t change).
  2. The document’s own source file, when it is itself an image — a standalone product photo ingested as one KB document (the primary shape for most fashion-ecommerce catalogs: one photo per SKU, not embedded in a PDF).
  3. Neither → a definite red cell, no provider call, no invented answer (R14).
Both sources resolve through StorageNamespace — the same tenant/project storage-namespace resolution every KB read uses — so a vision column can never cross a tenant’s storage boundary. Images are capped at KB_TABULAR_VISION_MAX_IMAGES (default 4) per cell — one catalog page with many figures never turns one cell’s generation into an unbounded provider call (SEC-LLM-001 gate 7). The provider call itself reuses the SdkAnonymousAgent + Base64Image pattern the OCR guide’s vision-llm driver already established — metered automatically by the laravel/ai SDK lifecycle hook, no double-counting, no new infrastructure. The flagship “Canonical KB Governance Audit” preset (#16 in the seeded library) is the headline application: rows = canonical docs, columns = 8 graph governance auditors + 1 verify contradiction check. Running it turns the whole KB into a per-cell-cited, exportable governance matrix — “which canonical docs are orphaned, stale, superseded, low-evidence-tier, or self-contradictory?” answered in one grid. The FE (built on the existing accessible DOM matrix) adds three surfaces: an agentic column editor (the governance metric picker appears only for graph columns; submit is gated so a graph-without-metric can’t 422), a per-cell evidence side-panel (summary + flag + reasoning + cited chunks), and a one-click template gallery of the built-in system reports.

Data model / contract

The engine reuses the v4.7 tables — no new tables for the agentic upgrade. The agentic dimension is additive on the column config: The ready-made library reuses workflows (type = tabular, is_system = true) — BuiltInWorkflowSeeder now mints 16 system templates (the +1 is the governance preset), each with a sensible columns_config. Both agent and metric are validated on Store/UpdateTabularReviewRequest (metric is required_if the agent is graph and must be one of the 10 GovernanceColumnResolver::METRICS). The capability is reachable on all three R44 surfaces over one shared core:
  • PHP — TabularReviewExtractor + GovernanceColumnResolver + VisionColumnResolver + the BuiltInWorkflowSeeder library.
  • HTTP — the existing api/admin/tabular-reviews/* (list / show / create / generate), now accepting agentic columns. vision needed zero new validation beyond the enum case — Rule::in(AgentKind::values()) already covered it.
  • MCP — App\Mcp\Tools\KbRunReportTool on the enterprise-kb server (read a saved report’s matrix, tenant-scoped R30, OFF-path safe R43), bounded by COUNT(DISTINCT) + LIMIT so an agent can’t pull an unbounded matrix. A vision column’s agent field passes through the same projection unchanged — confirmed by a dedicated test, not assumed.

Security & flags (R32 / R30 / R43)

  • RBAC. Every tabular-review route is behind can:viewTabularReviews / can:manageTabularReviews, R32-matrix-locked. The agentic upgrade adds no new route group — it extends the existing controllers’ request contracts.
  • Tenant isolation (R30). GovernanceColumnResolver scopes every kb_edges / document query to the active tenant; VisionColumnResolver resolves images exclusively through StorageNamespace::diskOf()/recordedPrefix() — the same namespace resolution every other KB storage read already goes through — so a vision cell can never read another tenant’s figures or source file; KbRunReportTool filters the report + cells by tenant. A client cannot read another tenant’s report or widen scope through a filter.
  • OFF-path safe (R43). KbRunReportTool returns a well-formed empty payload when the report/cells are absent rather than throwing; verify failures and graph unknowns degrade to a grey/red cell (R14) instead of erroring the whole generate. KB_TABULAR_VISION_ENABLED ships default OFF: a vision column with the flag off produces a definite red failed cell naming the flag — every other column in the same review still generates normally, never a 500, never a silent fallback to extract semantics.
  • Bounded provider work (SEC-LLM-001 gate 7). KB_TABULAR_VISION_MAX_IMAGES (default 4) caps how many images one cell’s provider call can carry — a catalog page with many figures cannot turn one cell into an unbounded multi-image call. The raw provider exception is Log::warning’d, never persisted into the cell’s reasoning (SEC-ERRLEAK-equivalent).

Decision rationale (ADR-style)

  • A per-column agent dimension, not a report-level mode. Governance audits mix deterministic facts and grounded judgments in the same grid; a whole-report mode would force a false choice. Making agent a column property lets one report carry graph facts next to extract/verify answers. See architecture decisions.
  • Deterministic graph metrics over LLM guesses. Canonical status, edge degree, orphan-ness and supersession are facts in the graph. Resolving them with GovernanceColumnResolver (no model call) is faster, free, and cannot hallucinate — and it reuses the real EvidenceTier/CanonicalStatus enums so the taxonomy never drifts (R9).
  • verify can only downgrade. An anti-hallucination pass that could raise a flag would itself be a hallucination risk. Constraining it to downgrade-only (and keeping the original cell on failure) makes the verified run monotonic and never worse than the extraction (R14).
  • Reuse workflows for the library. The ready-made “precotte” templates are just system-owned tabular workflows — no new model, no new admin surface, and they show up in the existing workflow gallery for free.
  • Glide canvas grid deferred. The accessible DOM matrix (per-cell testids + ARIA, R11/R15) ships now; the canvas grid + SSE progressive-paint are documented v8.19.x follow-ups because canvas cells can’t carry per-cell testids/ARIA — testability and a11y come first.
  • Image source is derived, not admin-configured. A vision column doesn’t ask the admin to point at a specific image — it derives the image(s) from the document itself (OCR figures, then a standalone image document). A per-column “pick an image” UI would be a materially bigger surface (an image picker, upload flow, multi-image gallery model) for a case the derivation already covers for both realistic shapes of a fashion-ecommerce catalog.
  • No manual cell-value override. “Human review in the grid” is satisfied by the existing review UX — flag colour, evidence panel, regenerate cell — the same loop every other agent kind already has. Editing a persisted cell’s value directly, without re-running extraction, would be a materially larger capability spanning every AgentKind, not a vision-specific one, and is explicitly deferred (ADR 0034 §6).

Worked example

Create a governance report from the flagship preset, then read its matrix over MCP. A graph column is computed with no LLM call:
A verify column downgrades a too-confident extraction:
Read the whole report through the MCP surface — an agent on the enterprise-kb server calls KbRunReportTool and gets the tenant’s report matrix, bounded by COUNT(DISTINCT) + LIMIT:
In the admin SPA at /app/admin/tabular-reviews, “From template” opens the gallery, picks “Canonical KB Governance Audit”, and pre-fills the create dialog; clicking a populated cell opens the evidence side-panel with the summary, flag, reasoning and the cited KB chunks. A vision column on a fashion catalog, KB_TABULAR_VISION_ENABLED=true:
The evidence panel renders that citation labelled “image” rather than “chunk” — the one FE change this feature needed (TabularReviewsList.tsx), since the value being cited is a filename the model was shown, not a retrieved chunk.

Gotchas & operations

  • graph needs a metric. A graph column without a governance metric is a misconfiguration — the FE gates submit on it and the BE validator rejects it (required_if), so a bad column can’t reach generation and 422.
  • graph wins over json_path. If a column sets both, the deterministic graph resolver runs; json_path is the LLM-free shortcut for extract-family columns.
  • verify adds a bounded second call. It is the one agent kind that costs extra per cell (one verify call after the extract). Use it for the columns where anti-hallucination matters, not blanket-on every column.
  • Existing reviews are unchanged. A pre-v8.19 review has no agent on its columns → every column is extract → byte-identical output. The agentic upgrade is purely additive (R27).
  • Re-run after canonical changes. graph cells reflect the canonical graph at generate time; promote/supersede/delete a doc and re-generate to refresh the governance matrix.
  • vision needs OCR figures or a standalone image document. A tenant using only the vision-llm OCR driver (which extracts no figures by design) and no standalone image ingestion gets an honest “no visual evidence” red cell for every vision column — never a fabricated answer. Switch to mistral or docling OCR (KB_OCR_DRIVER) to extract figures, or ingest product photos as standalone documents.
  • vision is one call per cell, not batched. A 100-row × 3-vision-column review generating from scratch issues up to 300 bounded provider calls, same shape as graph’s one-call-per-cell, not extract’s one-call-per-document batching. Real cost, bounded by KB_TABULAR_VISION_MAX_IMAGES and metered automatically.

Canonical & promotion

The canonical graph (kb_edges, status, evidence tiers) the graph columns read.

Grounding & evidence tiers

The EvidenceTier taxonomy a governance audit ranks each doc by.

Multi-tenant isolation

The TenantContext every report, cell and graph query is scoped to (R30).

MCP server

The KbRunReportTool reader on the enterprise-kb server.