Blog
Can AI read scanned documents and handwritten notes on a file?
Usually yes for clean scans, less reliably for poor ones, and rarely well for handwriting. What matters is whether the tool tells you which pages it could not read.
Alesis · · 5 min read
Most AI tools can read a clean scan as well as they read a typed document, because the words are recovered from the image before anything else happens. Poor scans, faxes of faxes, skewed pages and handwriting are a different matter: some of the text will come through wrong, and some will not come through at all. The question that matters for a firm is not whether the tool can read everything, but whether it tells you honestly what it could not read.
What happens when a scan is read
A PDF that came out of a word processor already contains text. A PDF that came out of a scanner contains a picture of a page. Before any question can be answered from that picture, the characters have to be recognised.
Recognition is good on modern flatbed scans of printed text. It gets worse with:
- pages photographed on a phone at an angle
- faint or heavily photocopied documents
- old faxes, including the header strip that runs across the top
- dense tables, columns and financial schedules where the layout carries meaning
- stamps, seals and signatures laid over printed text
- handwriting of any kind
When recognition fails it does not always fail loudly. A figure of 1,000 can come back as 1.000 or l,OOO. A date can lose a digit. A negative can lose the word "not" if the line is broken across a fold. These are quiet errors, and they are the ones to worry about, because nothing on the screen looks wrong.
Handwriting is the weak point
Treat handwritten material as unread until a person has read it. Attendance notes, margin annotations, a client's manuscript amendments to a draft, a doctor's records, a diary entry: these are exactly the documents that carry the point of the matter, and they are exactly the documents machine reading handles worst.
That does not mean an AI tool is useless on a file containing handwriting. It means the tool should be used for what it is good at, which is finding the printed material and telling you where the handwritten material sits, so that a fee earner can go and read those pages properly rather than hunting for them through six lever arch files.
The same applies to annotated printed documents. The typed text may be recovered perfectly while a manuscript "NOT AGREED" in the margin is missed entirely. If a page has been marked up by hand, the marked-up version is the one that matters, and a person should look at it.
The behaviour to insist on: flagged, not skipped
There is a real difference between a tool that skips a page it cannot read and a tool that flags it. A skipped page silently changes the answer. A flagged page tells you where to look.
When you are assessing a tool, put a deliberately bad document in front of it: a crooked phone photograph, a faint fax, a page of handwriting. Then ask a question the answer to which sits on that page. You want to see one of two things:
- The tool says it could not read that page, and names it.
- The tool says the papers do not answer the question.
What you do not want is a confident answer assembled from the pages it could read, with no mention of the pages it could not. Ask the supplier directly how unreadable pages are handled and whether the firm is told.
Citations have to point at pages
On a scanned bundle, an answer without a page reference is close to worthless, because you cannot check it without reading the whole thing yourself. If the tool names the page each point came from, checking a figure takes twenty seconds: open the page, look at the number, move on.
This is also the practical defence against quiet recognition errors. A misread date will survive inside a summary. It will not survive a fee earner glancing at the page it was taken from. Build that glance into how the work is done: any date, figure, name or limit that will be relied on gets checked against the page it is cited to.
Housekeeping that makes a real difference
Most of the improvement here is not technical. It is filing discipline.
- Scan at a sensible resolution, straight, in one pass, rather than page by page from a phone.
- Keep the original electronic version where one exists. Do not print a document and rescan it.
- Give files names that mean something, so a person can tell what they are looking at.
- Keep a single agreed pagination for a bundle and stick to it, so page references from any source line up.
- Separate correspondence from exhibits rather than merging everything into one enormous PDF.
- Flag on the file, in plain words, where the handwritten material lives.
A firm that does these things gets better results out of any tool, and better results out of its own people covering for each other.
Where Alesis fits
Alesis reads a matter's documents page by page, so its citations point at pages, and any page it could not read is flagged rather than skipped. It answers questions about a matter from the matter's own papers and names the page each answer came from; if the papers do not say, it says so. It assists qualified professionals and does not replace them, and it does not provide legal advice.