Observed Soundness: 96.3% of multimodal grading evaluations (233/242 live trials across 74 distinct prompts; 95% Wilson CI: [93.1%, 98.0%]) avoided materially incorrect guidance, spanning Pre-K through graduate work in mathematics, natural science, social science, language arts, foreign language, computer science, philosophy, law, and translation studies. Curriculum question quality holds at 95.8% (115/120 items; 95% Wilson CI: [90.6%, 98.2%]) across 12 graduate generation contexts.
Conservative Prompt-Cluster View: At the distinct-prompt level, 91.9% of prompts (68/74) were sound across every input mode tested (95% Wilson CI: [83.4%, 96.2%]).
Honest Distribution: Across 242 grading trials, 90.1% (218/242) were completely clean, 6.2% (15/242) had minor non-material phrasing nuances, and 3.7% (9/242) had flagged conceptual errors—all transparently audited below.
Adjudication Caveat: Every soundness figure reflects single-reviewer, non-blind adjudication (the two evaluation halves judged by different reviewers), not double-scored consensus. These measure the soundness of the guidance, not the displayed numeric score.
Anti-Slop Mechanism: Lune Synth eliminates unanchored scores via explicit confidence gates (τ = 0.85), deterministic rubric penalty catalogs, step deduction conflict checks, and 1-tap student verification to inspect and challenge every step.
Pipeline Execution: 100% transport and execution success across all 266 evaluated API and grading calls.
01 · Empirical Benchmark Summary (Side-by-Side Audit)
| Evaluation Dimension |
Sample Size (N) |
Clean (Zero Defect) |
Minor Nuance / Imprecision |
Material Failure |
95% Wilson Confidence Interval |
| Multimodal Feedback Soundness |
242 live responses (74 prompts, Pre-K–grad) |
90.1% (218 / 242) |
6.2% (15 / 242) |
3.7% (9 / 242) |
[93.1%, 98.0%] (Overall Soundness: 96.3%) |
| Prompt-Cluster Safety |
74 distinct prompts (Pre-K–grad) |
91.9% (68 / 74 sound) |
— |
8.1% (6 / 74) |
[83.4%, 96.2%] |
| Curriculum Item Quality (graduate-only) |
120 items (60 free, 60 choice) |
91.7% (110 / 120) |
4.2% (5 / 120) |
4.2% (5 / 120) |
[90.6%, 98.2%] (Defect-Free: 95.8%) |
| Complete 5-Item Set Soundness (graduate-only) |
24 full practice sets |
79.2% (19 / 24 pristine) |
— |
20.8% (5 / 24) |
[59.5%, 90.8%] |
| Pipeline Runtime Reliability |
266 evaluation API calls |
100.0% (266 / 266) |
0.0% drops |
0.0% fatal exceptions |
[98.6%, 100.0%] |
02 · Modality Stratification & High-Friction Stress Testing
Typed Digital Inputs (3 / 81 Material · 96.3% sound; 95% CI [89.7%, 98.7%]):
Tested across dense mathematical proofs, symbolic equations, code blocks, and multi-paragraph essays, from Pre-K arithmetic through graduate coursework.
Voice Audio Transcripts (2 / 81 Material · 97.5% sound; 95% CI [91.4%, 99.3%]):
Evaluated on conversational derivations, spoken step-by-step problem-solving, and spoken answers with natural pauses, revisions, and colloquial phrasing.
Handwriting & Scans (4 / 80 Material · 95.0% sound; 95% CI [87.8%, 98.0%]):
No modality effect is established: the Pre-K–high-school run recorded 0 material failures across 39 handwriting responses, and the combined handwriting interval overlaps both other modes. Caveat: 14 of those 39 trials ran as OCR-simulated text—fixtures quarantined by a 0.85 transcription-fidelity gate, including all five Pre-K cases—while 25 were genuine scanned images.
Distinct Prompt-Cluster Safety Rate (91.9%):
68 of 74 distinct prompt contexts avoided any material evaluation error across all student submissions in the cluster (95% Wilson CI: 83.4%–96.2%).
03 · Compound Set Quality & Defect Independence Model
Compounding Set Mathematics:
A question-level defect-free rate of p = 0.9583 (115/120) compounds across an independent 5-item mission set as:
P(Clean 5-Item Set) = p⁵ = (0.9583)⁵ ≈ 80.59%
This is consistent with the empirical benchmark result of 79.17% (19/24 full sets; 95% Wilson CI: [59.5%, 90.8%]), though with only 24 sets that interval spans a wide range of dependence structures and does not by itself establish independence. The observed 19/24 is the primary figure; the p⁵ value is a secondary projection.
04 · Evaluated Domains — Generation Contexts & Grading Coverage
Question-Generation Contexts (12 graduate disciplines):
The question-generation suite was benchmarked across 12 graduate-level academic and professional disciplines:
Complex Analysis
Immunology & Virology
Administrative Law
Compiler Optimization
Psychometrics (IRT)
Narratology & Literary Theory
Algebraic Topology
Quantum Information
Causal Inference
DSGE Macroeconomics
Historical Linguistics
Applied Cryptography
Grading Coverage (Pre-K through graduate): The 74 graded prompts spanned a broader set—mathematics, natural science, social science, language arts, foreign language, computer science, philosophy, law, literary analysis, translation studies, and qualitative methods—not limited to the 12 generation disciplines above.
05 · Anti-Slop Architecture & Deterministic Thresholds
Explicit Confidence Gates (τ = 0.85 Cutoff):
When optical character recognition or symbolic semantic parsing falls below strict certainty thresholds (τ < 0.85), confidence is immediately lowered and provisional review flags are surfaced.
Rubric Evidence Grounding:
Evaluations match against deterministic rubric criteria and explicit penalty catalogs rather than unconstrained holistic text generation, eliminating score drift.
Step-Level Conflict Checking:
Deduction arithmetic is programmatically reconciled against rubric point weights prior to finalizing step guidance, ensuring feedback and scores remain mathematically consistent.
1-Tap Student Agency & Verification Chain:
Students can inspect the full chain of evidence, contest ambiguous step deductions, challenge false negative classifications, and focus on the immediate corrective step.
06 · Defect Taxonomy & Full Failure Audit
Logged Failure Breakdown & Root Cause Analysis:
All 9 material grading errors (6 in the graduate tranche across 3 distinct prompts, plus 3 in the Pre-K–high-school run) and 5 question-generation defects were formally logged and cataloged:
- Grading failures — graduate tranche (6 responses across 3 distinct prompts): Four responses came from a single graduate astrophysics prompt on stellar fusion, where feedback endorsed thermal kinetic energy as classically overcoming the Coulomb barrier and treated quantum tunneling as optional detail; the same misconception recurred across typed, voice, and handwriting renderings of that one prompt. One response, on a philosophy-of-science scan, missed a polished but false conclusion that sufficient evidence yields a logically unique theory. One response, on a translation-studies scan, praised an incorrect tense/aspect analysis of the French imperfective retrouvais.
- Grading failures — Pre-K–high-school run (3 responses): Three material failures were logged in the 2026-08-19 evaluation and retained in the permanent corpus; per-case root-cause detail is tracked internally rather than on this page.
- Generation Defects (5 total): Contradictory Item Response Theory distractor rationale; non-unique option in Narratology; false-premise prompt in non-Hausdorff topology; underidentified fiscal shock equation in DSGE macroeconomics; and 1 ambiguous prompt boundary in Administrative Law.