AI can now describe a science diagram to a blind child
Question three on a science worksheet shows a cross-section of a plant stem, four parts labelled A to D, and asks which one carries water up from the roots. A screen reader will read that question aloud without any trouble at all. It will say nothing about the picture, which is where the whole answer lives.
For a child in Singapore who is blind or has low vision, the text of schoolwork stopped being the problem years ago. Screen readers handle it. Singapore’s official assistive technology guide, run by SG Enable, lists four things for visual impairment: braille devices, electronic magnifiers, screen readers such as NVDA and JAWS, and white canes. Some magnifiers include optical character recognition, which scans printed words and reads them aloud. All of that is for text.
Diagrams are different. They arrive either as a tactile version, which a human being has to make, or as a spoken description, which a human being has to give.
The bit of the worksheet that always needed a person
The Singapore Association of the Visually Handicapped runs a Braille Production Centre that transcribes English print into braille and produces what it calls “tactile diagrams and Braille signages”. It also offers a Braille Audit Service to check that transcriptions are accurate. That is careful, skilled work, and it is the reason a blind pupil can sit a paper with a diagram on it.
It is also work that happens on somebody else’s schedule. So is the help at Lighthouse School, which teaches children aged seven to eighteen with sensory impairment and staffs itself with teachers trained in braille. The Ministry of Education funds assistive technology devices for students with special educational needs in both mainstream and special education schools, and told Parliament in October 2024 that it “will continue to monitor technological developments, and adopt what is useful and appropriate to support the teaching and learning of all students, including students with SEN”.
None of that reaches a child at half past nine at night, stuck on question three, with a diagram in front of her and nobody awake to describe it.
The change is real, and you can see its size in the research
Point a phone at the page now and ask what the picture shows, and a multimodal model will tell you. The interesting thing is not that this works. It is how fast it got good, and there is a reasonably clean way to see that.
In 2024, a team at Cornell ran a two-week diary study with sixteen blind and low vision people using an AI scene description app in their ordinary lives. Participants rated the descriptions 2.76 out of 5 for satisfaction and 2.43 out of 4 for trust. The authors concluded that the descriptions “still need significant improvements to deliver satisfying and trustworthy experiences”.
Two years later, an overlapping team ran the same shape of study, two weeks, twenty blind and low vision participants, this time with an app built on a multimodal large language model. Satisfaction came in at 4.13 out of 5. These are separate studies with different apps rather than a controlled comparison, so treat the jump as an indication rather than a measurement. But it is a large one.
About one answer in five was wrong
The same 2026 study, accepted to the CHI conference, recorded something else. Of the responses participants received, 22.2 per cent were incorrect. The app declined to answer another 10.8 per cent of the time. Average trustworthiness rating: 3.76 out of 5.
So the honest version of the good news is this. A tool that was frustrating in 2024 is genuinely useful in 2026, and it is wrong roughly one time in five.
That error rate is not evenly spread, either. A survey of twenty peer-reviewed studies on describing STEM images, published in May 2026 by a team of accessibility researchers, found that factual inaccuracies and hallucinations remain a core unsolved problem for exactly the images school runs on: charts, diagrams, scientific figures. Their blunt summary was that “the research landscape is fragmented and its impact on real users remains limited”.
A sighted child can glance at the picture. That is the whole problem.
Every other child using AI for homework has a cheap error check available. They look. If the model says the graph peaks in March and the graph obviously peaks in July, they see it in a second.
A blind child has no such check. The description is not a hint about the picture, it is the picture. A confident wrong answer and a confident right answer sound identical.
This is why one piece of research from 2025 is worth more to a parent than any product announcement. Meng Chen, Akhil Iyer and Amy Pavel tried a simple idea: instead of showing one description, show several variations of it and let the disagreement between them do the work. With fifteen blind and low vision participants, the approach raised their ability to spot unreliable information by roughly 4.9 times. It also lowered how much they trusted the model, which was the point. Fourteen of the fifteen preferred it that way.
You do not need their prototype to use the finding. Ask twice, differently. “What does this diagram show?” and then “What are the labels, letter by letter?” If the two answers do not line up, something is off, and the child now knows it without needing eyes.
One caveat we should state plainly, because nobody else will: all of this research was done with adults. We could not find an equivalent diary study with blind children in a classroom, and that gap matters, because a ten-year-old is far less likely to push back on a machine that sounds certain.
Which makes the job here fairly specific, and quite different from the usual advice about limiting AI. Do not limit it. A child who can finally read a diagram alone at half past nine at night has gained something real. Just build the double-check in from day one, so asking twice feels like part of the tool rather than a sign of distrust. And keep the human line open for the things that have to be right: exam papers, assessment diagrams, anything that gets marked. Tactile production is slow because accuracy is the point.
If this is your child’s situation, our pieces on AI captions for deaf children in class and which AI reading tools actually help cover the same trade-off from different directions.
Was this useful?
Sources
- How Multimodal Large Language Models Support Access to Visual Information: A Diary Study With Blind and Low Vision People · arXiv (CHI 2026)
- Investigating Use Cases of AI-Powered Scene Description Applications for Blind and Low Vision People · arXiv (CHI 2024)
- Surfacing Variations to Calibrate Perceived Reliability of MLLM-generated Image Descriptions · arXiv
- A Systematic Survey on Image Description Techniques for STEM Domains · arXiv
- Enhancing Support for Special Needs Education and Potential Use of Assistive Technologies · Ministry of Education, Singapore
- Assistive Technology for Visual Impairment · Enabling Guide, SG Enable
- Braille Production Centre · Singapore Association of the Visually Handicapped