Product

AI grading for exams and essays, reviewed by the teacher

Closed questions grade themselves. On written answers, the AI suggests and you decide.

Evalmee grades the four closed question types automatically: true/false, single choice, MCQ and matching. The six open types stay marked by a human. On written answers, the Evalmee AI assistant suggests points and feedback, which you approve answer by answer before they count towards the grade.

What the Evalmee AI assistant does

The AI is never triggered on its own. You lock the grading rubric of the question, you click Generate feedback, you read it over, then you approve. Without approval, nothing enters the grade.

  • Manual trigger: one button, on the question you are grading
  • Suggestions, not grades: suggested points and written feedback
  • Explicit approval: grading is recorded only by your click
  • Anonymized data: no personal data is sent to the AI provider

Grading question by question, not submission by submission

The Grading page, By question tab, shows every participant's answer to the same question on a single page. You apply the same criterion forty times in a row instead of rebuilding it for each submission.

  • Every answer to a question on one page: the criterion stays stable
  • Less drift between the first and the last submission: that is the very point of this mode
  • Grading submission by submission: also available, from the Grades tab

Splitting the grading between several graders

Once your colleagues are invited to the exam, the split is made by question: you tick one or more questions, you click the grader's name, and they inherit those questions.

  • Assignment by question: tick, then click a grader
  • Three sharing roles: Full control, Editor, Read-only
  • Roles propagated: a role given on a folder applies to its subfolders and its exams

One example, from the rubric to the grade

Here is an assisted grading from start to finish, on a fictional open question: the rubric locked before the exam, a participant's answer, the Evalmee AI suggestion criterion by criterion, then the grader's decision. That decision, not the suggestion, is what enters the grade.

The question, out of 6 points: "Explain the difference between formative and summative assessment, with one example of each." Before the exam, the grader attached three criteria to the question, 2 points each, and locked the rubric: the definition of formative assessmentFormative assessmentAn assessment carried out during learning, whose function is to inform learner and teacher about what is currently being acquired. Its useful output is the feedback, not the mark: it is frequent, low-stakes, and serves to correct course while there is still time to do so. The term comes from Michael Scriven's work in 1967, taken up by Benjamin Bloom, which sets the assessment that forms against the assessment that takes stock. In practice a formative assessment is recognised by three traits: fast, precise feedback on every answer, a result that carries little or no weight in the final mark, and a teaching decision taken from the errors observed. Example: a ten-question quiz at the end of every session, marked automatically, with the reason for the correct answer shown at once; the teacher reads the per-question statistics and returns, in the next class, to the notion that half the group missed. The testing effect strengthens the case: testing yourself consolidates memory more than re-reading does. The same question can serve in a formative or a summative setting; what decides is the use made of the result, not the format., the definition of summative assessmentSummative assessmentAn assessment that comes at the end of a sequence, a module or a year to take stock of what has been learned and produce a mark that counts. It fixes a result rather than accompanying progress, and differs from formative assessment in its function, not in the format of the questions asked. The word comes from sum: the mark summarises what has been learned, and it serves to decide, to pass a module, to rank, to admit, to award a title. When that decision conditions a degree or a certification, the more precise term is certification assessment. Example: the final exam of a contract law course, two hours, one case study and ten short questions, marked out of 100 and weighing sixty per cent of the module mark, the rest coming from continuous assessment. Because its mark has consequences, a summative assessment demands more than a formative one: a marking scheme published in advance, identical sitting conditions for everyone, identification of the candidate, a traceable marking process and a route for appeals. An item analysis after the exam remains useful: a question that the strongest candidates miss more often than the others points to an ambiguous wording or a wrong answer key., and the relevance of the examples.

The participant's answer: "Formative assessment happens during learning. It is used to spot what the learner has not yet understood, so it can be revisited. Summative assessment is the one that counts towards the average, like the end-of-semester exam."

After a click on Generate feedback, Evalmee AI suggests 4 points out of 6 and this feedback: "The two definitions are told apart by their timing and by the use made of the mark. An example of formative assessment is missing; an ungraded end-of-class quiz would be one."

The grader reads it over. They keep the feedback and the point withheld on the examples, but restore the 2 points for the summative definition: "the one that counts towards the average" is the wording used in class. They click Approve. The answer is worth 5 points out of 6.

A fictional example. A three-criterion rubric, out of 6, as the question shows it to the grader.
CriterionPoints availableAI suggestionGrader's decision
Definition of formative assessment22: timing and purpose named2
Definition of summative assessment21: the purpose of taking stock of what was learnt is not named2: the wording is the one used in class
Relevance of the examples21: a single example, for summative1
Total645

How does AI-assisted grading work?

You define the criteria and the scoring before the exam. After the exam, you lock the rubric of the question, the AI generates suggested points and feedback, you read them over and you click Approve. The grade only moves at that last step.

The order matters: the rubricRubricA grid crossing assessment criteria with performance levels described in full sentences, from the weakest to the strongest. The marker places the paper in a cell rather than awarding a figure: the mark becomes the consequence of a description shared in advance. is locked before the AI steps in, so the scoringMarking schemeThe rule that converts a candidate's answers into a mark. It sets the number of points per question, any penalties, and the way partial credit is awarded. Published before the exam, it constitutes the examiner's commitment as to what will be rewarded. is settled before any submission has been seen and applies to every answer.

The assistant is called Evalmee AI. It relies on OpenAI, processing takes place in the European Union, and the data sent is anonymized. Client personal data is never shared with an AI provider, nor used to train models.

The documented course of an assisted grading, in criteria-based mode.
StepWho actsWhat happens
1. Assessment criteriaYouYou enter the criteria of the exam, one by one
2. Rubric by questionYouYou attach criteria to the question, write the grading instructions and set the points of each criterion
3. LockingYouThe rubric is frozen for every answer to the question
4. GenerationEvalmee AISuggested points and feedback, on explicit request
5. ReviewYouYou keep, adjust or replace the suggestions
6. ApprovalYouThe answer is graded, with the points assigned manually or by the AI

What does the AI do, and what does it not do?

The AI suggests points and writes feedback on written answers. It does not grade closed questions, defines no criteria, validates no submission and does not detect plagiarism. Each of these four limits matches a design choice, not a temporary gap.

A product page that promises fully automatic grading describes an exam few institutions would agree to defend before a board. So here is the exact line. The comparison of online exam platforms puts the same question to the six other French vendors, with the public page each answer was read on.

How the work is divided between automation, the AI assistant and the grader.
TaskWho does it
Grading a closed question: true/false, single choice, MCQ, matchingThe platform, deterministically. No AI is involved
Defining the criteria and the points of the rubricThe grader, before the exam
Suggesting points on a written answerEvalmee AI, on request
Writing feedback on a written answerEvalmee AI, on request, or the grader
Grading a spreadsheet, a diagram, code, an oral answer or an uploaded fileThe grader
Grading a scanned handwritten paperThe grader; the assistant works on answers typed into Evalmee
Validating the grading of an answerThe grader, by clicking Approve
Detecting plagiarism or AI-generated textTurnitin, a third-party service, with the client's licence
Proctoring the exam, detecting copy and paste and the second screenThe proctoring measures, without AI

Evalmee supports anonymous grading of submissionsAnonymous markingThe removal of the candidate's identity at marking time, the paper carrying nothing but a number. The practice targets biases linked to name, gender, presumed origin or a student's reputation, whose effect on the mark awarded is documented by research.. Two further fairness levers come on top: question-by-question grading, which reduces the severity gap between the first and the last submission, and locking the rubric, which prevents changing the scoring mid-grading.

The scoring scheme, before and after the exam

Before the exam, the scoring is set and the rubric locked. After the exam, it can still be changed from the Paper tab, and every grade is recalculated. In both cases, one exam is never graded under two scoring schemes.

A locked rubric cannot be changed without consequence: if you unlock it while answers have already been graded, that grading is reset. The paper follows the same rule. Once the exam is open, the questions seen by participants no longer change, so that everyone sat the same paper.

A scoring mistake can be fixed afterwards. From the Paper tab, the Edit button lets you revisit the correct choices, the weightings, the grading comment and the tags, then run Recalculate grades. A right answer ticked wrongly is fixed that way in two minutes, without reworking submissions one by one. A question to discard comes down to a weighting reset to zero, followed by the same recalculation.

On choice questions, the scoring is explicit and configurable, even after participants have sat the exam. The discordance mode, used in health-studies entrance exams, grades both the choices ticked and the choices left blank. To see what a penalty does to a paper before settling on one, the test grade calculator applies the same arithmetic question by question. To report the grade on a scale other than the point total, the score converter does the arithmetic and prints the full conversion chart.

The scoring settings of multiple-choice questions

  • Reward: all or nothing, or proportional to the number of correct choices
  • Penalty: none, cumulative or absolute
  • Calculation base: standard, balanced, moderate or manual
  • Floor: none, zero, or a value you set
  • Discordance scoringBarème par discordance (discordance scoring)A way of marking multiple-choice questions, common in the French PASS, LAS and former PACES health tracks, where the mark depends on the number of discordances with the expected answer: an item ticked wrongly, or a correct item left blank. Zero discordances scores full marks, and each further discordance drops the mark one tier. Each faculty sets the tiers itself.: the grade depends on the number of gaps between the participant's answer and the expected answer

Frequently asked questions

You define the criteria and the points before the exam, then you lock the rubric of the question. The Evalmee AI assistant then generates, on request, suggested points and feedback. You read them over and click Approve.

No. The four closed question types (true/false, single choice, MCQ, matching) are graded automatically, without AI. On written answers, the AI produces suggestions that the grader reads over then approves. It defines no criteria, validates no submission and does not detect plagiarism.

No. Evalmee AI works on answers typed into the Evalmee editor. A participant can attach a scan of their handwritten notes to an answer, but that file is graded by the grader, with no suggested points or feedback.

Grading assistance relies on OpenAI, with anonymized data and processing in the European Union. Client personal data is never shared with an AI provider, nor used to train models.

Yes, and it is the recommended mode. The Grading page, By question tab, shows every participant's answer to the same question. The aim is to reduce the severity gap between the first and the last submission graded.

Through Turnitin, which can be enabled in open-book mode and compares answers with online sources, with the answers of other participants and with AI-generated content. The service runs on your institution's Turnitin licence and only processes answers longer than twenty words.

Used by

  • Teaching teams
  • Exam boards
  • Freelance graders
  • Training centres

Grade your next session with the assistant

Start for free, no credit card required.