YakData

Production Case · Evaluation & Verification

The AI model is only one component.

77 Rules turns a dashboard and business context into a scored decision record a person can review. The case shows how image-and-text AI, explicit rules, evidence checks, fixed scoring logic, human control, and offline model testing divide responsibility so the model never owns the entire decision.

The key engineering decision: use the AI where visual interpretation helps, then surround it with rules and safeguards people can inspect, test, and own.

Evaluation & VerificationContext & DataOperational RiskDecision Ownership
AI perceivesIt does not own the final score.
Rules stay explicitExpert judgment remains inspectable.
Evidence controls claimsUncertainty changes system behavior.
Humans retain authorityReview and override are part of the workflow.

Interactive system map

Follow how the system runs, then inspect what sits underneath it.

Use the view controls to isolate the rule checks, evidence checks, or model test loop. Click any part to see what it is responsible for.

Interactive system map 77 Rules Internal System
HOW IT RUNS A dashboard becomes a scored decision record a person can review. 01 INPUTDashboard +business context 02 AI ANALYSISImage + textanalysis 03 RULE CHECKSExplicit rules77rulesplus additional integrity checks 04 EVIDENCEConfidence +evidence levelsA / B / C claim discipline 05 CONTROLFixed scoringrules100-point score + A to F grade 06 HUMAN REVIEWVerify +overrideEvidence remains explicit 07 OUTPUTDecision-ready review RULE GROUPS DESIGN CHECKSlayout · labeling · hierarchyvisual + communication quality CORE INTEGRITYdecision-risk checksevidence-aware logic ADDITIONAL INTEGRITYextended coveragesame controlled rule path EVIDENCE + CONTROL EVIDENCE LEVELSA directly visibleB needs verificationC context-dependent SCORING LOGICexplicit severity weightsfixed calculationfinal score + grade HUMAN CONTROLshow evidenceverify or overrideextend client rules FIX → EXTEND RULES → RE-RUN MODEL TEST LOOP Offline testing protects quality, reliability, speed, and cost per run. LABELED BENCHMARKKnown dashboard casesExpected findings as comparison answers CANDIDATE MODEL SETUPSReplaceable AI modelSame benchmark, comparable evidence MODEL TEST RESULTSquality | reliabilityspeed | cost PRODUCTION MODEL SETUPBest operating tradeoffquality · reliability · speed · costselection follows measured task performance
Drag to pan. Use + / − or the view controls to zoom. Click any node to inspect it.
Project File 001

The AI model is only one component.

The defensible system combines explicit expert rules, AI interpretation, confidence-aware claims, fixed scoring logic, human judgment, and a feedback loop that can absorb client-specific rules.

InputDashboard + context
Engine77 rules + integrity checks
ControlA / B / C confidence
OutputScored audit
What this page proves

Shown here: system design, rule structure, evidence handling, fixed scoring logic, human control, and the evaluation method. Client economics and model-provider comparison results are not part of this page.

Seven system responsibilities

Each layer owns a different kind of correctness.

The system works because AI analysis, expert rules, evidence checks, scoring, and human judgment are not collapsed into one giant prompt or one opaque AI response.

01

Input

Visible dashboard evidence enters with the business context required for claims that pixels alone cannot support.

02

AI analysis

The image-and-text AI interprets visual structure and returns evidence in a predictable format rather than acting as final authority.

03

Rule checks

Expert judgment becomes explicit checks that the software, tests, and users can inspect and understand.

04

Evidence checks

Confidence and evidence levels distinguish what is directly visible, what the AI is inferring, and what depends on supplied business context.

05

Fixed scoring control

Explicit scoring rules produce the final score and grade. The language model does not improvise it.

06

Human review

Users can inspect evidence, reject or correct findings, and extend the system with additional explicit rules.

07

Decision-ready output

The result is a prioritized, scored decision record that shows confidence and can be corrected and rerun instead of being a one-shot AI answer.

Claim discipline

If the pixels do not prove it, the system must not state it as fact.

A production AI system needs a clear way to separate what is visible from what the AI is inferring. In 77 Rules, confidence changes both the language shown to the user and what the system does with the finding.

A
Visually confirmable

The evidence is directly visible in the dashboard. The system can state the finding confidently.

B
Inferential concern

The pattern is suggestive but not fully proven. The system flags it for verification instead of presenting certainty.

C
Context-dependent

The finding fires only when supplied business context supports the claim.

Model testing

Choose the production model setup against the actual work.

The model is a replaceable part of the system. Candidate setups are tested against labeled cases, then compared on quality, repeatability, speed, and cost.

01Labeled benchmark

Known cases establish expected findings and the answers used for comparison.

02Candidate model setups

Replaceable model setups see the same benchmark and produce comparable evidence.

03Evaluation matrix

Quality, reliability, speed, and cost are considered together.

04Production tradeoff

Select the setup that clears the required task quality while balancing reliability, speed, and operating cost.

Lesson: optimize the operating system, not the prestige of the model inside it.

What to carry into your own build

Five design rules hidden inside the system.

01Validate AI output before your software uses it.

Do not let free-form AI text directly control what the app does.

02Keep consequential logic explicit.

Scoring, thresholds, permissions, and business rules should follow fixed rules where correctness requires it.

03Make uncertainty operational.

Confidence should alter wording, review behavior, and whether the system is allowed to make the claim.

04Put human control in the product.

Evidence, correction, rejection, and override belong in the workflow, not only in policy prose.

05Test models before production.

Choose the production model from labeled work, failure behavior, speed, and cost, not from how impressive it sounds in conversation.

A consequential AI decision in front of you?

Put the system question under scrutiny before the cost of being wrong compounds.

If the issue involves architecture, evaluation, economics, context, production readiness, operational risk, or decision ownership, YakData can pressure-test the system and determine what should happen next.

Discuss Your AI System Start With a Private Review Private Analysis Review: $2,500. Private Engagements: from $38,000.