Rules gather the facts and chase sign-off, AI summarises the year and drafts the text, and managers decide the rating and hold the conversation.
8 stepsTypical mix: AI candidateIllustrative analysisUpdated
Partly. AI can prepare most of a performance review, meaning it can gather the facts, summarise the year and draft the text, but a manager still has to decide the rating and hold the conversation. The process runs from collecting goals, results and feedback, through a summary and a written draft, to checking consistency, deciding the rating, having the conversation and filing the record. The slow part for most managers is the first half: finding the material and turning it into readable, specific text for each person. That is where a language model helps, because it reads long notes quickly and writes a clean first version. It does not know what happened in the room, it does not know that a good result last quarter came with a team conflict, and it cannot carry the responsibility for a judgement about a person. Treat it as a very fast assistant for the preparation and as a second pair of eyes for consistency. The decision, the tone and the accountability remain with people. If your reviews today are written the night before from memory, the first gain comes from keeping notes through the year, and AI makes that habit pay off.
In the UK the same process is often called an appraisal. If you are comparing performance review templates or software, look at how much of the preparation is covered and who can see the AI-assisted material, because the reviews are personal data under UK GDPR.
What each step needs
01● Standard automation
Collect goals, results and feedback for the review period
Before anyone writes a word, the facts have to be in one place: the goals set last time, the results against them, notes from one-to-ones, peer feedback, training completed, absence and any recognition. Today this is often a manager scrolling through old documents the night before. An integration or a simple scheduled export can pull these into one sheet per employee without any AI, because the sources are known and the fields are fixed. The point is that the manager starts from a complete file instead of from memory, which is where most recency bias comes from.
⛨Works when goals and one-to-one notes are kept in a system or shared document that can be exported. If notes live in private notebooks, the first job is agreeing where they go.
02● AI candidate
Summarise the period for each employee
A language model can read the collected notes, goal updates and feedback and produce a short neutral summary: what was achieved, which themes came up repeatedly, where feedback from different people agrees or differs. It can also quote the underlying notes so the manager can check each claim. This saves the hour of reading and sorting per person that makes reviews drag. The summary is a working document for the manager and not a verdict on the employee.
⛨Works when the input notes are factual and dated. Thin or vague notes give a thin, vague summary, and the model may smooth over gaps instead of pointing them out, so ask it to list what is missing.
03● AI candidate
Draft the written review text
From the summary and the manager's own bullet points, AI can draft the written review in the company's format and tone, with specific examples taken from the notes. Managers who find writing hard or who write twenty reviews in a week benefit most. The draft must be edited by the manager, who adds what they actually observed and removes anything they cannot stand behind. A review that reads as generic or that the employee recognises as machine-written damages trust faster than a short honest one.
⛨Works when the manager treats the draft as a first version and rewrites it in their own words. Do not paste unreviewed output into the system of record.
04● AI candidate
Check reviews for consistency and biased wording
Across a team or company, AI can compare reviews and flag patterns a busy HR team misses: one manager rating everyone as 'meets expectations', vague praise for some people and specific praise for others, personality words used for some groups and results words for others, or ratings that do not match the written text. It flags and explains; HR or the manager's manager decides what to do. This is a support step, because a flag can be wrong and the cure is a conversation, not an automatic correction.
⛨Works when the rating scale and criteria are the same for everyone being compared. Treat flags as questions to ask, and keep a record of what was changed and why.
05● Human review
Decide the rating and its consequences
The rating, and anything that follows from it such as a pay change, a bonus, a promotion or a performance plan, is a judgement about a person with real consequences. A manager makes it, a second person calibrates it, and the employee can challenge it. AI may give input, but a model should not set the rating, and in many places decisions with significant effects on people need human involvement by law. Calibration sessions, where managers compare cases, are one of the most useful hours in the whole process and they stay human.
⛨Always applies. Keep a short written reason behind each rating so that the decision can be explained to the employee and, if needed, to a tribunal or regulator.
06● Human review
Hold the review conversation
The meeting is where the review does its work: the employee hears what was seen, responds, disagrees, asks for help and agrees on goals. Tone, listening and the ability to handle a difficult message cannot be handed to software. AI can prepare the manager, for instance with a list of points to cover and questions to ask, but it should not be in the room as a substitute for the manager's attention. Employees who get a review they know was written by a tool, and read out by a manager, tend to disengage.
⛨Always applies. Give the manager the draft and the facts at least a few days before the meeting so there is time to think about how to say difficult things.
07● Standard automation
Send reminders, collect sign-off and file the record
The administration is a good fit for plain automation: opening the review cycle, nudging managers who are behind, sending the employee a self-assessment form, collecting signatures, storing the final document with the right access rights and creating the follow-up tasks from the agreed goals. A workflow tool or the HR system's own features do this reliably and cheaply, and no AI is needed. Doing it well removes the late chasing that makes managers dislike review season.
⛨Works when there is one agreed cycle with clear dates and an owner in HR. Access rights matter, because reviews are personal data that only the right people may read.
08● Keep human
Set the criteria, the scale and the rules for using AI
What 'good' means in each role, how many rating levels there are, how often you review, what the employee may see of AI-assisted material and what the AI may never be used for are management decisions. They also involve employee representatives in many countries and fall under data protection and, for AI that evaluates workers, rules that treat it as high risk in the EU. These are set by people, written down, explained to employees and revisited once a year.
⛨Always applies. Do this before the first AI-assisted cycle and tell employees plainly which steps use AI and who reads the result.
A sensible first experiment
Pick one team of eight to twelve people and one review cycle. Time how long a manager needs per review today, from collecting facts to finished text, and note how many reviews are late. In the pilot, let AI produce the summary and a first draft from the notes, while the manager edits everything and holds the conversation as usual. Measure the time per review, the share of reviews finished on schedule, and ask the employees in a short survey whether the review felt specific and fair. Keep the AI steps only if time falls and the employees rate fairness no lower than before.
The trap to avoid
The usual mistake is letting a tool draft the rating along with the text. The number then feels objective, managers stop questioning it, and any bias in the notes is repeated at scale. A second mistake is using AI on employee data before checking data protection and informing staff, which can undo trust and breach the law. Use AI for preparation and consistency checks, let people decide ratings, and tell employees what is used.
Questions teams ask
Can AI write performance reviews?
It can draft them from notes and feedback, and that is a real time saver. The manager must edit the draft, add what they saw first-hand and own the final text. A review written wholly by a tool and read out by a manager is easy to spot and weakens trust.
Should AI decide the performance rating?
No. The rating affects pay, promotion and sometimes employment, so it needs a human decision with a written reason, and calibration between managers. AI can help check that ratings are consistent, but it should not set them.
What does UK law say about using AI in appraisals?
Appraisals are personal data under UK GDPR, which also restricts decisions made solely by automated means when they have significant effects. The Equality Act 2010 applies to how ratings affect people. Keep a person in charge of the decision and tell staff how AI is used.
What do I need before using AI for reviews?
Written criteria, goals and notes that are kept in one place, a data protection check of the tool, and a clear message to employees about what is used and who reads it. Start with one team and keep the manager in charge of every review.
Illustrative workflow guidance by Arcgent. Each business needs its own assessment. No integration or savings claim has been verified for your systems.