Independent implementation · October 2026 · Review revision 2

Inspect a sales prediction, its meaning and its source

A numeric model estimates sales. A typed relationship graph supplies the metric definition and provenance. Code composes the main result; language-model output stays separate for inspection.

Built by Ingyu Koh with public advertising data. This demonstration makes model errors and review boundaries visible before connecting a workflow to business systems.

Validated inputs→Frozen prediction→Metric → model → dataset→Structured review boundary→Human decision

1. Try a scenario

Budgets: thousands of dollars. Sales: thousands of units.

Preset recordings load without model calls or generation quota. Custom inputs always receive a numeric prediction and a code-composed summary. Qwen’s idle CPU server is retired; Claude custom generation becomes available after AWS account activation. Live calls are bounded to 5/day per source network and 100/day globally.

Ready.

2. Inspect the result

—

Summary composed by code

Choose a scenario and run the prediction.

A human reviewer evaluates applicability. The prediction describes association; spending decisions need separate evidence.

Untrusted model output · diagnostic only
Choose “View model evidence” to inspect a recording.
Typed graph relationships and read-only tool trace

3. Test the review boundary

The previous keyword check accepted harmful spending instructions when they contained the expected number, unit and the word “causal.” Revision 2 removes that promotion path. Arbitrary model prose stays in the collapsed diagnostic panel; the main summary comes from trusted numeric tools and fixed text.

30/30Invalid or malicious probes rejected
10/10Valid structured controls accepted
20Actual Qwen generations inspected
0Claude results claimed before activation

The structured adapter permits seven fields with fixed choices, verifies them against tool results, and supplies numbers and units in code. This tests a narrow metadata task and its display boundary. The fixed probes measure this policy, rather than general language-model safety.

Full remediation evaluation · Original four-generation report, retained as history

Predictive evidence

1.976OLS test RMSE
5.829Mean baseline RMSE
40Held-out rows

RMSE is in thousands of sales units. The reproducible split uses 120 training, 40 validation and 40 test rows from 200 cross-sectional observations. Public textbook data demonstrates the integration; a client forecast needs its own time-based validation and agreed outcome.

Qwen/Qwen2.5-0.5B-Instruct has 494,032,768 frozen parameters. Its original free-text prompt and failed outputs are retained. Row 128 deliberately demonstrates the wrong-unit fallback. Cached inference timings describe the original generation, while browser timings describe retrieval.

Client implementation path

For the pet and garden business, Ingyu would agree sales and inventory definitions, validate a predictive baseline on authorized data, connect those definitions to the client’s semantic platform, and expose read-only tools for an agent. The delivery would include a test set, traceable outputs, operational limits and explicit human review steps.

This lab contains a small typed knowledge graph and two runnable MCP tools. The web path uses fixed read-only orchestration. The prepared Claude adapter uses Bedrock Converse with constrained JSON; runtime measurements await account activation. Databricks integration, an autonomous agent loop and production demand forecasting remain client-specific work.

Inspect and reproduce

Source, tests and deployment instructions · Measured AWS verification · Gateway health · Provider status · Data origin

Application logs record route, status, request ID and duration. Numeric inputs and model text are excluded. The 1 KiB body limit is checked after AWS receives the request; upload time is separate from handler execution.