The position
The mark is not the measurement
Indian schooling optimises hardest for the one measurement that cannot say what a student should do next. What follows is that problem, the harder problem underneath it, and the publication policy that both lead to: no methods, no positive results, and the findings that go against us published as they arrive.
- Published
- Positions, directions, open questions
- Not published
- Methods, parameters, results
- Written
- 29 August 2026
- Status
- Open to argument
Take two students who score sixty-two. One cannot handle a negative sign; the other cannot turn a sentence into an equation. The mark says the same thing about both, and says it in March.
The position in full
Four sections. The publication policy that follows from them is set apart below.
The mark is not the measurement
A mark is a scalar, produced at the end of a unit, a term or a board year, by a person with a stack of papers and a syllabus still to finish. It stands in for everything a student knows about a subject.
This is not an argument against examinations. A common paper, externally set and externally marked, is among the fairer instruments Indian schooling has: comparable across schools that share nothing else, portable, auditable, and indifferent to whom a family knows. It is also close to incorruptible, and that is the property any richer account of a student will find hardest to match. A detailed claim about a child, produced by a vendor, inside a system that allocates scarce seats, invites precisely the pressure the examination was built to resist.
The argument here is narrower. The mark is a serviceable instrument for ranking and a poor instrument for teaching, and Indian schooling uses it for both, because at the point where a decision about a child is made it has nothing else it trusts.
Consider two students who score sixty-two on the same mathematics paper. The first can carry out every procedure the paper asks for and loses marks in every long question to one habit: she drops the sign when she distributes a subtraction across a bracket. That is a single motor error, ten minutes of attention, correctable on a Tuesday. The second answered the quadratics section without a mistake and could not start the word problems, because he cannot turn four lines of English into an equation. His difficulty is not in mathematics. Both are sixty-two, both are entered in the register as average, and if the school runs remedial classes at all, both are sent to the same one and given more of the same paper.
What the compression discards is never general. It is always something specific, and always the thing the teacher would have needed: which topic, which misconception, which prerequisite, which skill. There is a fifth loss, different in kind from those four, and it is the one that should trouble people most: which moment.
Marks are produced after. After the unit is closed, after the class has moved on, after the misconception has been rehearsed for three weeks and become the student’s settled method. The window in which that sign error was cheap to fix opened in the second week and shut long before the report card was printed. A number that arrives after the moment in which it could have been used is not a measurement of learning. It is a record of an outcome.
Because the mark is the measurement the system trusts most, teaching bends towards it: previous years’ papers, the expected question, the marking scheme, the answer that scores. That is not a failure of teachers, who are behaving rationally under the only instrument they are held to. It is a property of the instrument. A measurement that compresses will, given enough pressure and enough years, produce a curriculum that compresses.
That argument does not stop at the mark. Anything proposed in its place inherits it, and a finer account of a child offers more surfaces to optimise than a single number does, not fewer. If a school is ranked on how many gaps it closes, gaps will close on paper. If a teacher is appraised on the same account, it becomes a document written for the appraiser rather than a description of a child. Our view is that this deformation would be worse than the mark’s, because a description is easier to curate than a common paper marked by a stranger, and the sign that it has begun is a school asking for the record to be improved rather than for the child to be taught. Before it buys anything, a school should ask any supplier proposing to stand in for the mark who sees the account, who is ranked on it, and what happens to it at an inspection.
No remedy is proposed here, and the reason is not modesty. Diagnostic alternatives to the mark have been put to Indian schools for a long time, some of them by the state. We have not surveyed that record and do not offer a survey of it here, so we name no programme and give no count. Our impression, which is an impression and not a finding, is that they failed on the teacher’s week more often than on their theory, by asking for time that does not exist. Any account of this problem that does not begin from the teacher’s week is describing a different country.
- Which topic
- A term aggregate of sixty-two can be a flat sixty-two across every chapter, or full marks everywhere and near-zero in one. The teacher cannot tell which, and the two demand opposite responses. Test: can the number distinguish a broad weakness from a single hole? It cannot.
- Which misconception
- Treating a function as though it distributes over addition is a specific, stable, teachable error with a name. In the mark it appears as three lost marks, indistinguishable from carelessness. The error is more diagnostic than the score, and it is the part that is thrown away.
- Which prerequisite
- A student who cannot balance a chemical equation is often not failing at chemistry but at ratio, four years earlier. Marked in chemistry, remediated in chemistry, and therefore not remediated. Prerequisite failures recur in new subjects, which is why they present as a general weakness.
- Which skill
- A physics answer marked wrong for units, a mathematics answer lost to arithmetic, a biology answer lost to the time taken to read the question in a second language. Three different failures, one column in the register. Language failure recorded as subject failure is, in our view, the most consequential misattribution in Indian assessment.
- Which moment
- The error was cheap on the Tuesday of the second week. The mark arrives in March. Everything between those two dates was rehearsal of the wrong method. Lateness is not a defect of implementation. It is what a summative instrument is for.
What academic intelligence means
Academic intelligence is not a product and not a polite synonym for artificial intelligence. Intelligence here carries the older sense it has in military intelligence or market intelligence: knowledge about a situation, assembled deliberately, in order to act well inside it. The adjective does the limiting work. Knowledge about learning, assembled in order to teach.
The distinction that matters is between a record and a model. A record states what a student produced: this answer, that score, this many minutes, on that date. Indian schools already hold far more records than they can use, and digitising a record does not change its kind. A model states what a record cannot: what a student knows, what she can do, and what she is ready for next.
Those are three different kinds of claim about a person, not one claim with three names. Knows is a claim about a present state that is not directly observable. Can do is a claim about a capability under conditions, and the conditions are doing more work than anyone admits. Is ready for is a claim about a future, inferred from a past that was thinly observed. A system that treats the three as interchangeable will be confidently wrong in a way that is hard to detect from outside it, and harder from inside.
This is why the difficulty is representational before it is technological. The questions below have to be answered before any question of method arises, and a claim that passes over them has not avoided answering them. It has answered them by accident.
We will not describe the representations we work with, here or anywhere on this site. Nor is that silence a settled answer being withheld: whether the useful unit is the student, a topic, or something narrower again is an open question here, and this site writes it as one. The point survives whatever anyone chooses to build: if the representation is wrong, nothing downstream repairs it, and a great deal downstream will make a wrong representation persuasive, which is to say fluent, confident, rendered in a dashboard and unfalsifiable. The field’s characteristic failure is not weak technology. It is strong technology resting on an unexamined account of what a learner is.
Academic intelligence, used as this page uses it, states an intent and not a capability. Nothing on this site should be read as a report that anything answering to it has been built, and anyone using the term, ourselves included, can fairly be asked which of the three claims above they mean and what observation would show them to be wrong about it.
- What is the unit
- A topic, a procedure and a misconception are three different kinds of thing, and they do not decompose into one another cleanly. Ask anyone who claims to know where a student stands which of them the claim is about. A direction that cannot name its unit of analysis has not yet stated a question.
- Whose property is knowing
- A child may execute a procedure on Monday and not on Thursday, in her first language and not her second, at the blackboard and not on paper, alone and not beside a friend. Knowing may be a property of a student and a context together, in which case a claim about the student alone is malformed. If that is so, the portability of a claim across contexts has to be demonstrated, not assumed.
- The same place
- Every grouping, every recommendation and every intervention rests on the phrase two students are at the same place. The phrase does the entire work and is almost never defined. Test: state the condition under which two students are at the same place, without referring to their scores.
- Ready for what
- Readiness is a claim about a future that has not happened, made from evidence gathered under conditions that will not recur. What licenses the word, and under what representation is such a claim even well formed? We do not have an answer to this we would defend in print.
Why this is a research problem and not a feature
A model of a learner is easy to make plausible and hard to make survive.
Plausibility is cheap because the case everyone reasons from is clean: one learner, dense and continuous evidence, an attentive tutor who can ask a follow-up question the moment something looks wrong, a single language, unlimited time, no syllabus deadline. Much of what we have read about adaptive learning describes, structurally, that situation. The Indian classroom inverts nearly every one of those conditions at once, and the inversions interact.
What is genuinely unknown is whether an account of a learner that is adequate under those conditions degrades gracefully or collapses. Our working belief is that it collapses, and that the collapse is not accuracy falling by some margin but the model’s claims quietly changing meaning. A claim earned in a tutoring setting becomes, in a room of fifty-five with shared devices and circulated notebooks, a claim about something else entirely, while continuing to be worded in the same way.
There is also an evidence problem that no care on our part resolves. We have not conducted a systematic review of the literature on learning technology and do not offer one here, so this site states no proportion and makes no claim about that literature as a whole. The premise we work from is narrower: the studies we have actually read were run in other class sizes, other languages, other assessment cultures and other incentive structures, and transfer to a Class X section in a Tier-2 Indian school is something a supplier should have to show rather than assume. That is a question a school can put to anyone, us included. Where was your evidence gathered, in what class size, in which language of instruction, and who paid for it.
One blunt test, which we would want applied to any claim of this kind and have applied to none of ours. Show an experienced teacher several accounts a system has produced of students in her own section, with the names removed, and mix in among them an account of a child she has never taught. Ask her to place them. If she cannot do better than she would by guessing, or cannot reject the stranger, she has been shown fluent prose and not an account of anyone. If she can, two answers are worth having, and the more useful of the two is that is wrong about Divya, because a falsification is something a person can act on the same afternoon. A test with no foil in it measures plausibility, which is the property such prose has whether or not there is anything behind it.
What we do not know, stated plainly. Whether a representation adequate for procedural mathematics carries over to a subject in which the answer is an argument rather than a result. How much of what presents as a knowledge gap in an English-medium classroom is a language gap wearing its clothes; we have not measured that fraction, and if it is large, a great deal of what is read as evidence about what a student knows is evidence about how fast she reads a second language. Whether the evidence a real school can afford to collect, in the time it actually has, is sufficient in principle to support the claims people want to make from it. And the uncomfortable one: whether a teacher told something true and specific about a student on Tuesday can act on it at all, or whether time binds so tightly that better information changes nothing. If that is the answer, most of this work is beside the point, and we have no basis yet for assuming it is not.
- Evidence is sparse
- A teacher across five sections may see any individual student’s unaided written work once a fortnight. Inference about a person from that much signal is a hard statistical situation, not an engineering inconvenience. Sparsity is structural. Instrumenting more of the room does not produce more independent work from any one child.
- Evidence is not independent
- Students sit close, help one another, and circulate the strong student’s notebook the night before submission. This is ordinary classroom life, not misconduct, and it breaks an assumption that is easy to make without noticing: that a piece of work is evidence about the person who submitted it. Treating shared work as independent evidence produces a model that is most confident exactly where it is most wrong.
- A wrong answer is ambiguous
- The same blank space can mean the concept was never understood, the question was misread, the English was too slow, the time ran out, or the student had decided by then that the paper was lost. Disambiguating these requires evidence the classroom does not naturally produce. A system that guesses between them fails invisibly, because the output reads the same either way.
- Identity is uncertain
- Devices are shared between siblings, between friends, and with a parent. An interaction record is not reliably a record of one person. Any claim about a student rests on an identity assumption that is often simply untrue, and the error compounds silently over a term.
- The calendar does not move
- The syllabus must be finished. A system whose recommendation is to spend three more days on this chapter is recommending something no teacher in the country is permitted to do. Usefulness is bounded by what is actionable inside a fixed calendar, which rules out any advice that assumes the calendar is negotiable.
- The spread is inside one room
- A single section commonly contains students years apart in a prerequisite. Whole-class recommendations are therefore wrong for most of the class by construction. This is the condition that makes the problem interesting, and the one a school should ask any supplier whether its evidence was gathered under.
What would make this worth doing
The standard is narrow, it is inconvenient, and it is the one we would want applied to anybody else making claims about Indian classrooms.
The work is worth doing if a teacher of fifty-five, given no additional hours in her week, ends a Tuesday knowing one true and specific thing about one named student that she would otherwise have learnt in March, and can do something about it before Thursday. That is the whole of the ambition. It is a smaller claim than the field usually makes and considerably harder to satisfy, because every constraint in it is real: no extra time, a named student, an actionable window, and true.
We have not met that standard, and nothing on this page should be read as a report that we have. The conditions below are what we would accept as failure, stated in advance rather than discovered later in a testimonial.
- It must be able to fail
- There has to be an observation that would tell us the work is not worth continuing. If no evidence could do that, we are not doing research, whatever the output looks like. Applies to each direction separately. A programme that survives every possible result is a marketing plan.
- The gain must land at the bottom of the room
- The strongest students in any section have always found ways to learn, and almost anything will help them. If the average improves because the already-capable improved, the thing worth doing has not been done. The distributional question has to be asked before the average one, or the average will hide the answer.
- The teacher must be able to overrule it
- When a teacher contradicts an account of her student, she is usually right, and the disagreement is the most informative thing in the exchange. It has to be recorded and treated as evidence, not absorbed as noise. A system that cannot be contradicted in a way that changes it is not being evaluated by anyone.
- It must hold in the second language
- For a large share of Indian students the medium of instruction is not the language of the home. Anything that works only when comprehension is not the binding constraint has been evaluated on the easy half of the problem. Language failure recorded as subject failure is the specific error to guard against.
- It must respect what the calendar allows
- Advice a teacher cannot act on inside a fixed syllabus is not advice. Usefulness is bounded by what fits into the week the school actually has. Advice that assumes the calendar is negotiable is ruled out before anything else about it is considered.
- A school must be able to leave
- A school that stops working with us should be able to take what it has accumulated and go. Dependence is not evidence of value. If the case for staying rests on the cost of leaving, the work did not earn the position it holds.
Disclosure · the licence this site runs under
What we publish, and what we do not
This site publishes positions, research directions and open questions. It does not publish methods, parameters, architectures, system descriptions or results.
Results are withheld because a result without its method attached is not a result. A percentage in a paragraph, an effect size in a brochure, an accuracy figure beside a logo: these have the grammar of evidence and none of its substance, and a reader has no way to interrogate them. The harm is not mainly that one buyer is misled. It is that a school leader who reads enough of them stops being able to tell any two claims apart, and the market then rewards whoever prints the largest number. If we ever hold a result worth having, it will appear as a study, with its method stated, its setting described, its sample and exclusions declared, its analysis pre-specified, and someone without our interest in the outcome involved in producing it. The route for that runs through the Council’s research instrument, which was published as a founding draft on 29 August 2026 and is open for comment. The Council has no members, no ethics review panel has been appointed, and no study has begun under it. It is a commitment we have written down, not a protection a school can rely on today, and nothing has yet gone through it.
Methods are withheld for a plainer reason we see no need to dress up. Work that has not been published as a study will not be leaked in outline through a diagram on a research page. Anything sufficient for a competent engineer to reconstruct is, for our purposes, publication, and we would rather publish when the work can be judged than while it can only be admired.
One class of finding is not withheld, and the asymmetry is deliberate. Every direction on this site names an observation that would show it to be wrong. When such an observation arrives, the outcome is published here as a bare direction, abandoned or revised or sustained, with the date it changed, whether or not a study is ever written about it. A null finding is published on the same terms. A finding that a direction is right is not claimed at all until it can be published as a study with its method attached. The half of the record that could be used to sell something is the half that has to wait; the half that costs us is published as it arrives, and a reader deciding whether any of this is load-bearing should watch the status of the directions rather than this paragraph.
What remains after the two exclusions is the contestable part: a position a reader can disagree with, a problem described precisely enough that a rival account of it can be given, a question stated with its unit of analysis named, and a statement of what would show us to be wrong. That material can be argued with by someone who has never seen our systems, which is exactly the property we want it to have.
One disclosure belongs in the text rather than a footnote. PrepGraph sells into this market. Any finding of ours that happens to favour our own products should be discounted until a party without our interest in the outcome has reproduced it. The Council for Responsible AI and School Innovation is voluntary and non-statutory, convened and funded by PrepGraph; it has no members as yet, it certifies nothing, and it endorses no vendor, PrepGraph included. A body a vendor funds cannot confer legitimacy on that vendor, and is not asked to.
Ledger What leaves this site, and what does not
Published
- Positions
- Research directions
- Open questions
Not published
- Methods
- Parameters
- Architectures
- System descriptions
- Results
Published as it arrives
- A finding that a direction is wrong
- A null finding
Not claimed
- A finding that a direction is right, until it can be published as a study with its method attached
Colophon
Notes on this page
- This site publishes no platform figures by design: no counts of students, schools, items or sessions, and no accuracy, benchmark or effect-size numbers. What is published instead, and what a reader can hold this site to, is the status of each research direction and the date on which it last changed.
- This page was written and published on 29 August 2026 and has not been revised since. Every position on it carries that date. When one is withdrawn or changed, the change and its date are recorded on this site rather than made silently.
- No study has begun, and nothing here has been through external review.
- Disclosure: PrepGraph is a commercial vendor in the market this research concerns. Readers should discount any finding of ours that favours our own products until it has been reproduced by a party without our interest in the outcome.
- The PrepGraph Council for Responsible AI and School Innovation, India, is a voluntary, non-statutory forum convened and funded by PrepGraph. It certifies nothing and endorses no vendor, including PrepGraph. Its research instrument was published as a founding draft on 29 August 2026 and is open for comment; the Council has no members, no ethics review panel has been appointed, and no study has begun under it.
- Related properties: the commercial learning platform at prepgraph.com, teacher professional development at academy.prepgraph.com, the schools and partners portal at partners.prepgraph.com, and the Council at council.prepgraph.com.
- Classroom conditions referred to throughout, including sections of forty to sixty and the fifty-five used as a running example, describe the ordinary Indian secondary classroom. They are not measurements of anything we operate.
Next
Where this argument is taken up
Research
Four Research Directions
These four directions were chosen because each names something that is broken in an ordinary Indian classroom and that we do not know how to fix.
Read the four directionsPrinciples
Ten Positions on AI in School
Ten positions on what AI should and should not do to learning. Each is stated so that a competent person could reject it, and each carries a demand a head teacher could make of any supplier, ourselves included.
Read the ten positions