Input
Visible dashboard evidence enters with the business context required for claims that pixels alone cannot support.
Production Case · Evaluation & Verification
77 Rules turns a dashboard and business context into a scored decision record a person can review. The case shows how image-and-text AI, explicit rules, evidence checks, fixed scoring logic, human control, and offline model testing divide responsibility so the model never owns the entire decision.
The key engineering decision: use the AI where visual interpretation helps, then surround it with rules and safeguards people can inspect, test, and own.
Interactive system map
Use the view controls to isolate the rule checks, evidence checks, or model test loop. Click any part to see what it is responsible for.
The defensible system combines explicit expert rules, AI interpretation, confidence-aware claims, fixed scoring logic, human judgment, and a feedback loop that can absorb client-specific rules.
Shown here: system design, rule structure, evidence handling, fixed scoring logic, human control, and the evaluation method. Client economics and model-provider comparison results are not part of this page.
Seven system responsibilities
The system works because AI analysis, expert rules, evidence checks, scoring, and human judgment are not collapsed into one giant prompt or one opaque AI response.
Visible dashboard evidence enters with the business context required for claims that pixels alone cannot support.
The image-and-text AI interprets visual structure and returns evidence in a predictable format rather than acting as final authority.
Expert judgment becomes explicit checks that the software, tests, and users can inspect and understand.
Confidence and evidence levels distinguish what is directly visible, what the AI is inferring, and what depends on supplied business context.
Explicit scoring rules produce the final score and grade. The language model does not improvise it.
Users can inspect evidence, reject or correct findings, and extend the system with additional explicit rules.
The result is a prioritized, scored decision record that shows confidence and can be corrected and rerun instead of being a one-shot AI answer.
Claim discipline
A production AI system needs a clear way to separate what is visible from what the AI is inferring. In 77 Rules, confidence changes both the language shown to the user and what the system does with the finding.
The evidence is directly visible in the dashboard. The system can state the finding confidently.
The pattern is suggestive but not fully proven. The system flags it for verification instead of presenting certainty.
The finding fires only when supplied business context supports the claim.
Model testing
The model is a replaceable part of the system. Candidate setups are tested against labeled cases, then compared on quality, repeatability, speed, and cost.
Known cases establish expected findings and the answers used for comparison.
Replaceable model setups see the same benchmark and produce comparable evidence.
Quality, reliability, speed, and cost are considered together.
Select the setup that clears the required task quality while balancing reliability, speed, and operating cost.
Lesson: optimize the operating system, not the prestige of the model inside it.
What to carry into your own build
Do not let free-form AI text directly control what the app does.
Scoring, thresholds, permissions, and business rules should follow fixed rules where correctness requires it.
Confidence should alter wording, review behavior, and whether the system is allowed to make the claim.
Evidence, correction, rejection, and override belong in the workflow, not only in policy prose.
Choose the production model from labeled work, failure behavior, speed, and cost, not from how impressive it sounds in conversation.
A consequential AI decision in front of you?
If the issue involves architecture, evaluation, economics, context, production readiness, operational risk, or decision ownership, YakData can pressure-test the system and determine what should happen next.