Blog
How do we know if AI is actually saving our firm time?
Measure at task level, not firm level. Pick a few repeated jobs, record how long they took before, then compare like with like including the time spent checking the output.
Alesis · · 5 min read
You find out by measuring a small number of repeated tasks, not the firm as a whole. Record how long a task took before, run it the new way for a few weeks, and compare the two honestly, counting every minute spent checking and correcting the output. Anything broader than that tends to produce a number nobody trusts and nobody acts on.
Measure tasks, not people or the firm
Firm-wide productivity figures are almost useless for this. Too many other things move at the same time: case mix, staff changes, a quiet August, one unusually heavy matter. If chargeable hours go up or down after you introduce a tool, you will not be able to say why.
Task-level comparison is different. Pick three or four jobs that happen often enough to compare, are reasonably similar each time, and have a clear start and end. Good candidates tend to be:
- Building a chronology from a new file.
- Summarising a bundle of correspondence for a client update.
- Finding and reading the current version of a provision or a piece of official guidance.
- Producing a first draft of a routine letter or schedule.
- Working out a date calculation and setting out the reasoning.
Bad candidates are the one-off, the unusual and the strategic. You will never get a clean comparison on advisory work that depends on a partner's judgement.
Get a baseline before you change anything
This is the step most firms skip, and it is the step that makes the exercise worth doing. For two or three weeks before you start, ask the people doing the chosen tasks to note the time each one took. Not an estimate written up later: a note made at the time.
Keep the baseline crude but consistent. Task name, fee earner, matter type, minutes, and a one-line note on anything unusual. Ten or fifteen data points per task is enough to see a pattern. You are not running a study; you are trying to avoid arguing about numbers later.
If your time recording is granular enough to pull this from existing data, use it. Most narratives are too vague, which is why a short manual exercise usually beats mining the system.
Count the checking time, or the number is fiction
A draft that arrives in two minutes and takes forty minutes to verify has not saved forty-eight minutes. It may still be a gain, but only if you compare it against what the same work took from scratch.
So record the whole cycle:
- Time framing the question or assembling the papers.
- Time waiting for and reading the output.
- Time checking it against the source: the page, the section, the guidance.
- Time correcting, redoing or abandoning it.
The fourth item matters most. Attempts that went nowhere are part of the cost. A tool that works well four times in five and wastes an hour on the fifth is a different proposition from one that works well every time, and only a full record will show you which you have.
Be alert to work that simply moves. If a junior now produces a draft in half the time but a partner spends longer reviewing it because they trust it less, the firm has saved nothing. Ask reviewers whether their time on the same task has changed.
Measure quality and risk as well as minutes
Time saved is the easiest thing to count and the least interesting on its own. Alongside the minutes, keep a short record of:
- Errors found in checking. How many, how serious, and whether they were obvious or subtle. Subtle errors that a busy fee earner might miss are the ones that should worry you.
- Whether the tool said when it did not know. A tool that flags a gap is safer than one that fills it smoothly.
- Whether the source was traceable. If you cannot get from a statement to the page or the provision it came from, checking will always be slow.
- Work you would not otherwise have done. Sometimes the gain is not speed but coverage: reading the whole of a document set rather than a sample. That is a quality improvement, not a time saving, and worth recording as such.
You may also find that a task gets no faster but gets less unpleasant, and that the fee earner does it earlier rather than putting it off. That is a real benefit, though you should be honest that it is not the same as a time saving.
Decide in advance what you will do with the answer
Agree the decision before you collect the data. For example: if the chronology task comes in meaningfully faster with no increase in errors found at review, it becomes standard practice and goes in the supervision notes. If it does not, you stop using it for that task and say why.
Set a review date and put a named person on it, whoever holds responsibility for AI in the firm. Look at the cost side at the same point: what the firm actually spent over the period, against the measured effect. Then repeat the exercise for the next few tasks rather than assuming the result carries across. A tool that is strong at reading a file may be weak at drafting, and the only way to know is to measure each one.
Where Alesis fits
Alesis reads a matter's papers page by page and names the page each answer came from, and it names the source for every point it makes on legislation, official guidance and Financial Ombudsman decisions, which is what makes the checking step quick enough to measure. When it cannot find support for a point, it says what is missing rather than guessing. It is funded by credit rather than a subscription, so the cost side of any comparison is simply what the firm chose to spend. It prepares drafts for a qualified person to review and sign off, and it does not provide legal advice.