How do we know this?
Knowledge explains. Trust reassures. Evidence proves. This page exists because every important claim on this website should be able to answer one question — and a claim that cannot is not finished.
Si Math AI is a comprehensive learning platform for SAT, ACT, and EST Mathematics that combines educational expertise, AI technology, personalized learning, analytics, and human support to help students improve their understanding and performance.
Artificial Intelligence is how Si Math AI teaches.
Educational expertise is what it teaches.
Human experience is why it works.
Three kinds of evidence — and one we do not have
Edtech marketing routinely blurs different kinds of support into one undifferentiated impression of rigour. Everything below is labelled, so you can weigh it yourself:
- MechanismWhat the software demonstrably does. Verifiable by using the product — the strongest evidence here, because you can check it rather than trust us.
- ResearchAn established finding in the educational literature that supports the principle a feature applies. Note carefully what this does and does not mean — see below.
- RecordA dated, documented artefact in our repository — a migration, an incident record, an audit, a regression test that pins a fix in place.
The kind we do not have: outcome evidence. We have no measured results showing that Si Math AI's own students improve more than students using something else. Producing that honestly requires a controlled study and a large sample, and we have neither. Until then, no page on this site will imply otherwise — and you should treat any platform claiming outcome evidence without describing how it was collected with the same scepticism.
What citing research does and does not mean. Citing Roediger & Karpicke does not mean research shows Si Math AI works. It means retrieval practice works, and Si Math AI applies retrieval practice. Conflating those two is the most common dishonesty in this industry. Where the literature complicates a claim we benefit from, we say so — the caveats below are not decoration.
The methodology is the product
Si Math AI is an educational methodology implemented through software. The software delivers the methodology; it is not the methodology itself.
Software can be copied. An educational philosophy cannot. That is why this page begins with the method rather than the features: the method is the claim, so the method is what needs evidence attached to it.
Si Math AI is not built around Artificial Intelligence. It is built around Educational Intelligence. Artificial Intelligence is simply one of the tools used to deliver that educational intelligence.
Students do not improve because they use AI. Students improve because they follow a better learning process. AI simply makes that learning process scalable, personalized, and available between lessons.
It follows that technology alone improves nothing. It is worth something only in combination with educational expertise, sound teaching methodology, meaningful practice, continuous feedback. Without those elements, AI becomes just another chatbot. With them, it becomes an educational accelerator.
The 8 components, and what supports each. Note which one is listed last — artificial intelligence is the delivery mechanism, not the advantage:
Expert Mathematics Teaching
The curriculum, explanations, mistake catalogue and exam strategy are authored and reviewed by experienced SAT, ACT and EST mathematics educators. This is the foundation the other seven rest on.
No citation claimed
Continuous Personalized Assessment
Every attempt — in chat, a drill or a mock — is assessed against one canonical skill rather than left as undifferentiated activity, so assessment is continuous instead of episodic.
Weakness Analysis
Assessment is aggregated into a ranked diagnosis: which specific skills are costing the most marks, with a severity band each. Students are poor judges of their own weaknesses, which is why this is external rather than self-reported.
Evidence-Based Revision
What a student revises is decided by their own recorded evidence, not by chapter order or by what feels uncomfortable.
Deliberate Practice
Practice is targeted at identified weaknesses and skips skills already mastered — effortful and specific rather than high-volume and comfortable.
Long-Term Knowledge Retention
The method is built for what survives to test day, not for what a student can do in the lesson. Spacing and interleaving are the mechanism; Learning Memory is what makes them possible across months.
Human Educational Experience
Judgements about what to teach and how to teach it stay with people. Specialists review how students actually perform and revise the material accordingly.
No citation claimed
AI-Assisted Personalization
Artificial intelligence delivers the seven components above to one student at a time, at any hour, in the language they think in. It is the delivery mechanism, listed last because that is where it belongs.
Two components carry no citation, and that is deliberate. No paper supports "our teachers are experienced". Attaching one would be the exact conflation this page warns about elsewhere, so the gap is shown rather than filled.
Evidence for every capability
For each of the 8 major systems: what it is, why it exists, how it works, what supports its value — and, deliberately, the limits of that support.
Zero AI Mentor
The tutoring layer. Zero works through a problem step by step, offers a different approach when the first does not land, reads photographed questions, and explains why the wrong answer choices are wrong.
Expert explanation is the scarcest resource in exam preparation. One teacher with thirty students cannot give each of them an unhurried, patient walk-through of the question they are stuck on tonight. Zero exists to remove the scheduling constraint on expert explanation — not to replace the expertise, which is authored by educators.
Zero delivers teaching methods, mistake diagnoses and exam strategies that human specialists authored and review. Every interaction also writes a diagnostic signal against a specific skill, which is what feeds the rest of the platform. A scope guard declines requests outside the platform's educational purpose.
- MechanismExplanations state why distractors are wrong, not only why the key is right — checkable in any session. Distractors on standardised exams are engineered around specific misconceptions, so naming the misconception is the teaching act.
- ResearchThe worked-example effect: for learners without strong prior knowledge, studying fully worked solutions is more efficient than problem-solving alone, because unguided search consumes the working memory that learning requires. See the research →
- ResearchFeedback is among the most powerful influences on achievement — but its effect varies enormously with type. Feedback about the task and the process helps; feedback directed at the self does not. Zero is built to comment on the work, never on the student. See the research →
- RecordA scope guard is implemented in the ai-tutor Edge Function and pinned by a regression test (tests/scope-guardrail.test.mjs), so out-of-scope requests are declined and no diagnostic record is written for them.
The limits of this evidence: We have no measured outcome data showing that students taught by Zero improve more than students taught another way. We are not going to imply otherwise. What we can show is that the mechanism follows established practice and that the content is specialist-authored.
Weakness Analyzer
Converts every attempt into a diagnostic signal against a specific skill, then ranks weak skills by how much each is costing the student's score, with a severity band per skill.
Students are poor judges of their own weaknesses. The topics that feel hardest are frequently not the ones losing marks, and errors like sign-handling under time pressure are close to invisible from the inside. Without an outside view, study time goes to the wrong skills — which is the single most common reason preparation stalls.
Signals from chats, focus drills and mock exams resolve to one permanent skill identifier in a fixed taxonomy before storage, so "quadratics", "quadratic equations" and "solving by factoring" cannot fragment into three unrelated records. Ranking combines how heavily a skill is tested with how reliably the student loses it. A weakness clears only when later evidence confirms the fix held.
- MechanismThe taxonomy is a fixed, versioned artefact in the codebase (5 topic domains, 33 skills, permanent identifiers). Every stored record carries the taxonomy version that produced it, so a curriculum revision cannot silently rewrite what an old diagnosis meant.
- ResearchLearners' judgments of their own learning are systematically miscalibrated — fluency during study is mistaken for durable knowledge. This is the direct empirical case for an external diagnosis rather than self-assessment. See the research →
- ResearchFormative assessment — using assessment to direct subsequent teaching rather than only to grade — is associated with substantial learning gains. The Weakness Analyzer is formative assessment applied continuously rather than at intervals. See the research →
- RecordSeverity bands, recency weighting, trend, and question-level linkage were added in dated migrations on 2026-06-14; taxonomy v1 was applied across every consuming table in nine migrations from 2026-06-25 to 2026-06-28.
The limits of this evidence: The ranking is a model, and models are wrong sometimes. A skill can be ranked as weak because of two unlucky attempts. This is why the platform re-tests rather than closing a weakness on a single correct answer — but it does mean an individual ranking should be read as evidence, not verdict.
Focus Practice
Generates targeted drill sets from the ranked weaknesses, in the order that recovers the most marks per hour, skipping skills already mastered.
A diagnosis with no action attached is just a more precise way of feeling bad about your preparation. Focus Practice is the step that converts diagnosis into work — and the step that stops students practising what is comfortable, which is the most common way study hours are wasted.
Reads the live weakness ranking and the time remaining before the exam date. Every drill attempt writes a new signal, so the practice that fixes a weakness is also the practice that proves it was fixed. The queue re-ranks as weaknesses clear.
- MechanismPractice is generated from the student's own diagnosis rather than a fixed chapter order — two students the same distance from the same exam receive different sets. Observable by comparing two accounts.
- ResearchA large review of learning techniques rated practice testing and distributed practice as high-utility, and rereading and highlighting as low-utility. Focus Practice is practice testing, scheduled. See the research →
- ResearchInterleaved practice of mathematics problems produces worse performance during practice and better performance on later tests than blocked practice — because blocked practice never trains the hardest step, deciding which method applies. See the research →
- ResearchDesirable difficulties: conditions that slow acquisition often improve retention and transfer. This is why effective practice feels worse than ineffective practice, and why a platform optimising for how productive a session feels would optimise the wrong thing. See the research →
- RecordSignal-aware focus plans, plan lifecycle and single-active-plan enforcement shipped in dated migrations on 2026-06-15 and 2026-06-16; XP for completed focus work on 2026-06-24.
The limits of this evidence: Focus Practice can only target skills the taxonomy represents. A weakness that falls outside the 33 tracked skills will not be surfaced by it, and the taxonomy is deliberately conservative rather than exhaustive.
Mock Exams
Full-length, correctly timed SAT, ACT and EST mathematics mocks. You sit the paper under real timing, then record your result and the mistakes you made.
Practising questions and sitting an exam are different skills. Under real timing, students who know the mathematics still lose marks to pacing, fatigue, triage and the decision of when to abandon a question. A mock is the only way to measure that and the only way to practise it.
Real format and timing per exam, the result recorded by the student rather than converted by the platform, mistakes retained for review, and every recorded mistake fed back into the Weakness Analyzer.
- MechanismMocks are full-length and timed to each exam's real duration rather than sampled, and the mistakes you record are carried into the Weakness Analyzer rather than left in the session — both directly observable. The platform does not convert your raw count into a scaled score; you record the score your own result shows.
- ResearchThe testing effect: retrieving information from memory produces more durable learning than restudying it for the same time. A mock exam is retrieval practice at maximum scale, which is why it is a learning event and not only a measurement. See the research →
- ResearchMathematics anxiety consumes working memory, degrading performance independently of knowledge. Rehearsing real conditions is the mechanism by which test-day novelty — and the anxiety that comes with it — is reduced. See the research →
- RecordExam session idempotency was added on 2026-06-16, so a network interruption cannot double-record or corrupt a submitted attempt.
The limits of this evidence: Mock exams are our own construction, not retired official papers. They follow the published format and are specialist-reviewed, but a mock score is an estimate of exam performance rather than an exam result — and the platform performs no score conversion of its own.
Smart Progress Tracking
Mastery score per skill, trend lines over time, a predicted test-day score, daily streaks and a seven-rank XP ladder.
The most demotivating property of self-directed preparation is invisibility: you study for six weeks with no reliable way to know whether it worked. Measurement makes improvement legible — and stagnation legible early enough to change what you are doing.
Mastery is derived from actual attempts rather than questions completed. Trends are computed per topic domain. The predicted score updates as evidence accumulates. Streaks and ranks sit on top as motivational structure, never as a substitute for the mastery scores.
- MechanismMastery is computed from attempt outcomes, not from activity volume — a student can complete many questions without mastery moving, which is exactly the signal a volume metric would hide.
- RecordThe rank ladder exists in exactly one place in the client (assets/ranks.js), is mirrored by the SQL function rank_for_xp(), and a drift test (tests/constants-drift.test.mjs) fails the build if any copy diverges — so a student cannot see a different rank on different pages.
- RecordStreak columns and their recomputation shipped 2026-06-16; the streak logic is covered by its own regression suite (tests/streak.test.mjs) after an audit found it seeded today unconditionally.
- ResearchBloom's mastery-learning work argued that individual tutoring with formative feedback substantially outperforms conventional instruction. We cite it as the origin of the aspiration and note plainly that the headline two-sigma effect has proven difficult to replicate at that magnitude. See the research →
The limits of this evidence: The predicted test-day score is an estimate produced by our model from platform performance. It is not an official score, not a guarantee, and it has not been validated against a large sample of real exam results — because we do not yet have one. Read the direction, not the digit.
Learning Memory
A persistent, searchable record of every question, session and mistake, carrying context between sessions.
Exam preparation is cumulative. A tutor whose memory resets between sessions cannot do the one thing a tutor is for — notice a pattern across time. Every other system on the platform depends on this one: without persistence there is no ranking, no trend, no predicted score and no plan.
Sessions are stored and re-openable, history is searchable, mistake records persist as structured data, and exam context carries forward.
- MechanismOpen a session from three weeks ago and it is still there, with its explanation intact. This is the single clearest observable difference from a stateless assistant, and it takes one minute to verify.
- ResearchThe spacing effect — distributed practice outperforms massed practice for the same total time — is one of the most robust findings in the study of memory. Acting on it requires knowing what a student studied and when, which requires persistence. See the research →
- RecordQuestion-record idempotency (2026-06-16) ensures a retried write cannot duplicate a student's history, which is what makes longitudinal counts trustworthy rather than approximately right.
The limits of this evidence: Persistence is also a responsibility. This is student data, held to power that student's own diagnosis; it is protected by row-level security on every table and can be permanently deleted by the student. See the Trust Center for the full account.
Personalized Learning
The adaptive layer: what a student practises, in what order, how a concept is explained, in which language, and for which exam — all derived from evidence about that student.
A generic plan is not a neutral default. It systematically wastes the time of every student it does not happen to fit, and it hides the specific gaps that are actually costing marks.
Selection and sequencing are driven by the live diagnosis and the time remaining before the exam. The curriculum itself is fixed and specialist-authored — the pedagogy is chosen per student, not improvised per student.
- MechanismThe curriculum/selection split is the checkable part: the material a student receives comes from a specialist-authored body of content, and what changes per student is which of it they get and when. This is why the platform can answer "what is my child being taught?" with a stable answer.
- ResearchThe formative-assessment literature is about adjusting instruction in response to evidence of learning. Personalization here means exactly that, and nothing more mystical. See the research →
- ResearchWe deliberately do NOT personalize by "learning style" — the evidence for matching instruction to visual/auditory/kinaesthetic preferences does not support the practice, despite its popularity. Adapting to a student's measured weaknesses is well supported; adapting to a self-reported style is not. See the research →
The limits of this evidence: That last point is worth stating loudly because it costs us a marketing line. "Personalized to your learning style" is a claim many platforms make and the evidence does not support it. We personalize to diagnosed weakness, exam and language — things that can be measured.
Human Support
Educators and exam specialists author and review everything the platform teaches; people handle accounts and upgrades; specialists revise material based on how students perform.
Educational technology fails in two predictable directions: built by engineers without teachers it produces impressive software that does not reflect how learning works; built by teachers without engineers it produces sound pedagogy that cannot scale or measure itself.
Specialists define and maintain the skill taxonomy, author and review explanation methods and strategy content, and review student performance to revise material. Upgrade requests are reviewed and activated by a person with email confirmation within 24 hours.
- MechanismThe taxonomy is a hand-authored, versioned artefact with permanent identifiers and a documented amendment process — the visible trace of specialist curriculum work rather than model output.
- RecordPayment approval is admin-gated and takes the credit amount server-side from the plan definition rather than from the client request — a human decision, implemented so it cannot be forged.
- ResearchFeedback quality — not feedback quantity — determines effect. Deciding which explanation to give a student with a particular misconception is a teaching judgement, which is why specialists author the material rather than approving whatever a model produced. See the research →
The limits of this evidence: Human support does not currently include on-demand live tutoring. If that is what a student needs, a private tutor alongside the platform is the right combination and we say so on the Trust Center.
Research foundations
Si Math AI did not discover any of this, and says so. These are established findings in the study of learning, and the platform's contribution is applying them consistently to SAT, ACT and EST mathematics preparation — not inventing the principles.
References are given as author, year and title so you can look the work up directly. We deliberately do not quote effect sizes: summarised second-hand they are frequently misleading, and the papers are more useful than our gloss on them.
The spacing effect
Practice distributed over time produces more durable retention than the same amount of practice massed together. Among the most replicated findings in the study of human memory, first documented in the nineteenth century and confirmed across a very large modern literature.
How Si Math AI applies it: Why Si Math AI is built around short regular sessions, why streaks exist, and why previously weak skills are revisited rather than assumed fixed.
- Ebbinghaus, H. (1885). Über das Gedächtnis (Memory: A Contribution to Experimental Psychology).
- Cepeda, N. J., Pashler, H., Vul, E., Wixted, J. T., & Rohrer, D. (2006). Distributed practice in verbal recall tasks: A review and quantitative synthesis. Psychological Bulletin.
The testing effect (retrieval practice)
Retrieving information from memory strengthens it more than restudying the same material for the same time. Reading a solution feels like learning and largely is not; attempting the problem first is what consolidates.
How Si Math AI applies it: Why the guides tell students to attempt before checking, why Focus Practice is drills rather than reading, and why mock exams are treated as learning events rather than only measurements.
- Roediger, H. L., & Karpicke, J. D. (2006). Test-enhanced learning: Taking memory tests improves long-term retention. Psychological Science.
- Karpicke, J. D., & Roediger, H. L. (2008). The critical importance of retrieval for learning. Science.
Not all study techniques are equal
A large review assessed common study techniques for effectiveness. Practice testing and distributed practice were rated high utility. Rereading, highlighting and summarisation — the three most popular techniques among students — were rated low utility.
How Si Math AI applies it: The direct basis for the study-strategies guide, and for the platform not offering features that make ineffective study feel productive.
- Dunlosky, J., Rawson, K. A., Marsh, E. J., Nathan, M. J., & Willingham, D. T. (2013). Improving students' learning with effective learning techniques. Psychological Science in the Public Interest.
Interleaving beats blocking in mathematics
Mixing problem types within a practice session produces worse performance during practice and better performance on delayed tests than practising one type at a time. Blocked practice tells the student which method applies before they read the question — so it never trains the step that actually matters on an exam.
How Si Math AI applies it: Why Focus Practice mixes topics rather than grouping them, and why the guides recommend blocking only while first learning something new.
- Rohrer, D., & Taylor, K. (2007). The shuffling of mathematics problems improves learning. Instructional Science.
Desirable difficulties
Conditions that make learning feel harder and slower during acquisition often improve long-term retention and transfer. A corollary that matters commercially: the study methods that feel most productive are frequently the least effective.
How Si Math AI applies it: Why the platform optimises for measured mastery rather than for how satisfying a session feels, and why the guides warn students that effective practice feels worse.
- Bjork, R. A. (1994). Memory and metamemory considerations in the training of human beings. In Metcognition: Knowing about Knowing.
Feedback works — but its type decides whether it helps
Feedback is among the strongest influences on achievement, but effects vary widely and can be negative. Feedback about the task, the process and self-regulation tends to help; feedback directed at the self ("you are clever", "you are bad at this") tends not to, and can harm.
How Si Math AI applies it: Why Zero comments on the work and the method rather than on the student, and why the platform frames a mistake as information about a skill rather than as a verdict on a person.
- Hattie, J., & Timperley, H. (2007). The power of feedback. Review of Educational Research.
- Kluger, A. N., & DeNisi, A. (1996). The effects of feedback interventions on performance. Psychological Bulletin.
Caveat: This is a literature that complicates as much as it supports. It is cited here because it constrains our design, not because it endorses our product.
Mathematics anxiety consumes working memory
Anxiety occupies the same limited working-memory resources that multi-step mathematics requires. This is the mechanism behind a student who knows the material and still underperforms under pressure — the knowledge is intact; the capacity to deploy it is reduced.
How Si Math AI applies it: The basis for the exam-psychology guide, for treating rehearsal under real conditions as the primary intervention, and for the platform's position that pressure applied by parents tends to reduce performance rather than increase it.
- Ashcraft, M. H., & Kirk, E. P. (2001). The relationships among working memory, math anxiety, and performance. Journal of Experimental Psychology: General.
- Beilock, S. L., & Carr, T. H. (2005). When high-powered people fail: Working memory and "choking under pressure" in math. Psychological Science.
Cognitive load and the worked-example effect
Working memory is narrowly limited. For learners without strong prior knowledge, studying worked examples is more efficient than unguided problem-solving, because search consumes the capacity that learning needs. The advantage reverses as expertise grows.
How Si Math AI applies it: Why explanations are worked step by step rather than compressed, why the guides tell students to write intermediate steps down rather than hold them mentally, and why the platform moves a student from worked explanation toward independent practice as mastery rises.
- Sweller, J. (1988). Cognitive load during problem solving: Effects on learning. Cognitive Science.
- Sweller, J., & Cooper, G. A. (1985). The use of worked examples as a substitute for problem solving in learning algebra. Cognition and Instruction.
Students misjudge their own learning
Learners' judgments of what they have learned are systematically miscalibrated. Fluency during study — material feeling easy because it is in front of you — is routinely mistaken for durable knowledge, and students consequently stop studying material they have not yet learned.
How Si Math AI applies it: The empirical case for the Weakness Analyzer. If self-assessment were reliable, an external diagnosis would be unnecessary; it is not, so it is.
- Kornell, N., & Bjork, R. A. (2007). The promise and perils of self-regulated study. Psychonomic Bulletin & Review.
Formative assessment improves learning
Assessment used to direct what happens next — rather than only to grade what already happened — is associated with substantial gains, particularly for lower-attaining students.
How Si Math AI applies it: The whole diagnose → target → re-measure loop. The Weakness Analyzer is formative assessment running continuously instead of at intervals.
- Black, P., & Wiliam, D. (1998). Inside the black box: Raising standards through classroom assessment.
- Black, P., & Wiliam, D. (1998). Assessment and classroom learning. Assessment in Education.
Mastery learning and individual tutoring
Bloom reported that students taught individually with mastery-learning methods substantially outperformed conventionally taught students, and framed the challenge of achieving comparable results at scale as an open problem.
How Si Math AI applies it: The origin of the platform's aspiration: deliver something closer to individual, diagnosis-driven instruction to students who cannot access a personal tutor.
- Bloom, B. S. (1984). The 2 sigma problem: The search for methods of group instruction as effective as one-to-one tutoring. Educational Researcher.
Caveat: Stated honestly: the headline two-sigma magnitude has proven difficult to replicate, and later work suggests a smaller effect. We cite Bloom as the source of the question, not as evidence for our answer to it. Any platform quoting "two sigma" as a projected benefit of its product is overreaching.
Learning styles are not supported by the evidence
The popular idea that instruction should be matched to a student's visual, auditory or kinaesthetic "learning style" lacks credible supporting evidence, despite its very wide adoption.
How Si Math AI applies it: Why Si Math AI does not personalize by learning style — a deliberate omission, not an oversight. We adapt to diagnosed weakness, exam and language, all of which are measurable.
- Pashler, H., McDaniel, M., Rohrer, D., & Bjork, R. (2008). Learning styles: Concepts and evidence. Psychological Science in the Public Interest.
How features are built
Deliberately, and with a record. The pattern below is visible in the repository rather than aspirational — each stage leaves an artefact you can point at.
-
Research and design
A capability starts as a written design or architecture note stating what problem it solves and what it will not do. Educators define the educational requirement; engineers define the mechanism.
-
Engineering review before it ships
Non-trivial work gets a written review — including a migration risk assessment where the database is involved. Reviews are kept whether or not they were flattering.
-
Verification, not vibes
Behaviour is pinned by regression tests that fail the build if it changes. An internal audit of our own verification framework specifically hunted for checks that passed vacuously — a test that cannot fail is worse than no test, because it produces false confidence.
-
Release gate and closeout
Significant milestones pass a release gate and get a closeout document reconciling what the repository says with what production actually does. Documentation that has drifted from production is a defect.
-
Student feedback and specialist review
Specialists review how students actually perform and revise the material. Educational content is reviewed rather than shipped on model output alone — a plausible explanation is not automatically the right one to teach.
-
Incidents are recorded
When something breaks, an incident record is written and kept. One appears in the changelog, because a changelog containing only successes is a marketing document.
Evidence published elsewhere on this site
Architecture
The full learning flow, stage by stage, and why each stage exists — including the design decisions and their trade-offs.
Changelog
Dated development history, every entry traceable to an artefact in the repository — including an incident.
Roadmap
Direction rather than promises, led by the gaps we have publicly admitted to — and what we will deliberately never build.
Trust Center
What the platform does not do, when a human teacher is better, how data is protected, and why no testimonials are published.
Free educational guides
Twelve guides teaching the method in full, free and without sign-up. The most direct evidence of teaching expertise available.
Educational principles
Six positions, each stating not only what we believe but how it changed the software.