Assessment glossary

The vocabulary of assessment, defined term by term: question formats, marking schemes, psychometrics, exam proctoring, interoperability standards, accessibility and the French regulatory framework. Every definition stands on its own — it can be read and quoted in isolation.

85 terms defined. Regulatory and standards terms link through to the body that owns the definition; none of them is defined from memory.

Question and answer formats

What the candidate is asked to produce — and what each format can genuinely measure.

MCQ (multiple-choice question)

A closed question offering several possible answers, one or more of which are correct; the candidate ticks the ones they accept. The format allows instant automatic marking and covers a lot of content in little time, but measures poorly the ability to produce, structure or justify an argument.

Create an MCQ with EvalmeeBarème par discordance (discordance scoring)

Single-choice question

A variant of the multiple-choice question in which exactly one of the options offered is correct. The candidate can tick only one item, which simplifies the marking scheme and removes partial-credit rules, but increases the role of chance: four options give a twenty-five per cent chance of guessing right.

Short open-answer question

A question calling for a written answer of a few words to a few lines, with no options offered. It removes guessing and forces the candidate to recall rather than recognise, while staying short enough to be marked quickly, either by hand or with automated assistance.

Essay question

A question that calls for a developed argument: an essay, a case study, a proof. It is the only format that measures argument, structure and the ability to draw on several areas of knowledge at once, at the cost of slow marking and a marker-to-marker variability that a rubric exists to reduce.

Mark written answers

True/false question

A closed question with two possible answers. Quick to write and quick to mark, it leaves a fifty per cent chance of guessing right: it is only reliable in series, across many items, and lends itself poorly to assessing nuanced reasoning or partial knowledge.

Matching question

A question asking the candidate to link the elements of two lists: a term and its definition, a date and an event, a molecule and its function. It tests command of a vocabulary or a nomenclature efficiently, and is marked automatically, pair by pair.

Fill-in-the-blanks (cloze) question

A passage of text from which certain words have been removed and which the candidate must restore, either freely or from a list supplied with the question. The format sits between multiple choice and open answer: it demands active recall while still remaining automatically markable.

Item

The elementary unit of a test: a question and the marking scheme attached to it. This is the level at which the statistics of an exam are computed — difficulty, discrimination, correlation with the total score — and therefore the level at which an exam is repaired, question by question.

Distractor

An incorrect option in a multiple-choice question. A good distractor is plausible to anyone who has not understood and clearly ruled out by anyone who has: a distractor nobody ever chooses carries no information and reduces the question to a narrower choice.

Question bank

A reservoir of indexed, reusable questions, classified by topic, level or competency. It allows an exam to be assembled by random draw, papers to vary from one session to the next, and the statistical history that makes a question reliable to accumulate, item by item.

Marking schemes, grading and marking

How answers become a mark — and what guarantees that two markers reach the same result.

Marking scheme

The rule that converts a candidate's answers into a mark. It sets the number of points per question, any penalties, and the way partial credit is awarded. Published before the exam, it constitutes the examiner's commitment as to what will be rewarded.

Barème par discordance (discordance scoring)

Criterion-referenced marking scheme

A marking scheme that distributes points across explicit criteria — accuracy, method, writing, justification — rather than on an overall impression of the paper. Each criterion is marked separately, which makes the final mark decomposable, explainable to the candidate and comparable from one marker to another.

Barème par discordance (discordance scoring)

A way of marking multiple-choice questions, common in the French PASS, LAS and former PACES health tracks, where the mark depends on the number of discordances with the expected answer: an item ticked wrongly, or a correct item left blank. Zero discordances scores full marks, and each further discordance drops the mark one tier. Each faculty sets the tiers itself.

Discordance scoring in Evalmee

Rubric

A grid crossing assessment criteria with performance levels described in full sentences, from the weakest to the strongest. The marker places the paper in a cell rather than awarding a figure: the mark becomes the consequence of a description shared in advance.

Automatic marking

The award of points by the platform, with no human intervention, for questions whose expected answer can be decided in advance: multiple choice, matching, fill-in-the-blanks, a numeric value within a stated tolerance. It makes the result immediate and removes marker-to-marker variability on those items entirely.

AI-assisted marking

Pre-marking of written answers by a language model, which proposes a mark and a justification that the marker validates, amends or rejects. The decision stays human: what the AI saves is the first pass over the paper, never the responsibility for the mark.

AI-assisted marking at Evalmee

Anonymous marking

The removal of the candidate's identity at marking time, the paper carrying nothing but a number. The practice targets biases linked to name, gender, presumed origin or a student's reputation, whose effect on the mark awarded is documented by research.

Double marking

The marking of the same paper by two independent examiners, whose marks are then compared. The gap between the two measures how much of the exam is subjective; beyond a threshold agreed in advance, a third reading or a discussion between markers settles it.

Moderation

A meeting of markers, before or after marking, to align their reading of the marking scheme on sample papers. It reduces the systematic gap between markers — the fact that the same piece of work is worth two points more in one pile of papers than in another.

Floor mark

A guaranteed minimum mark, below which a paper is not taken however weak it is. It limits the destructive effect of a zero on a yearly average and neutralises accumulated penalties, at the cost of compressing the marking scale towards the top.

Negative marking

The deduction of points for wrong answers to a closed question, intended to discourage answering at random. The practice also penalises reasoned risk-taking and, according to the research available on the subject, disadvantages cautious candidates more than it disadvantages the ones who are poorly prepared.

Barème par discordance (discordance scoring)

Compensation

The mechanism by which a mark above the pass line in one subject offsets a mark below it in another, within a teaching unit or a semester. Its scope, its threshold and the subjects excluded from it are set by the institution's assessment regulations.

Resit session

A second examination session, open to candidates who did not meet the progression requirements at the end of the first. Its timetable and the rule on whether the original mark is kept or replaced are set by the institution's assessment regulations.

Psychometrics and item analysis

The indicators that tell you whether an exam really measures anything, and which of its questions are defective.

Item analysis

Statistical examination of the questions in a test after it has been sat: how many candidates got each item right, which ones did, and how that compares with their overall score. It reveals ambiguous questions, useless questions and answer-key errors before they are repeated.

Per-question statistics

Difficulty index

The proportion of candidates answering an item correctly, between zero and one. An index of 0.90 flags a question almost everyone gets right, and therefore one that carries little information; values close to 0.5 maximise a question's ability to separate levels of attainment.

Discrimination index

A measure of a question's ability to tell strong candidates from weak ones, computed as the gap in success rates between the highest and the lowest overall scorers. An index of zero or below flags a defective item: badly worded, or with a wrong answer key.

Item-total correlation

The correlation between success on a question and the score obtained on the rest of the test, often expressed as a point-biserial coefficient. It shows whether the item measures the same thing as the exam as a whole; a low value flags an item that is off-topic.

Reliability

The consistency of a measurement: how far the same candidate would obtain a comparable score on sitting the test again, with other questions of the same kind or in front of a different marker. An unreliable exam produces marks part of which is nothing but noise.

Cronbach's alpha

A coefficient summarising the internal consistency of a test, that is, how far its items measure a single dimension. It generally varies between zero and one and rises mechanically with the number of questions: a high alpha therefore does not on its own attest to validity.

Validity

The degree to which a test really measures what it claims to measure and supports the decisions drawn from it. An exam can be perfectly reliable and have no validity: it then measures something very stable, but not the competency it claims to assess.

Content validity

The fit between the questions asked and the domain the exam is supposed to cover. It is established before the exam is sat, by comparing the spread of items against the syllabus or the competency framework: an exam covering three chapters out of twelve has no content validity.

Standard error of measurement

The margin of uncertainty attached to a candidate's score, expressed in the units of the mark. It is a reminder that eleven and twelve may not be distinguishable, and an invitation not to base a pass-or-fail decision on a gap smaller than that margin.

Standardised score

A mark brought onto a common scale, usually by subtracting the group mean and then dividing by the standard deviation. It allows candidates assessed on different papers or in different sessions to be compared, by placing each one relative to the distribution of their own group.

Differential item functioning

A situation in which, at equal levels of competency, candidates from different groups do not have the same probability of succeeding on an item. Differential functioning flags a potential bias — cultural, linguistic, gendered — lodged in the wording of the question rather than in the competency assessed.

Purposes and timing of assessment

The same question does not say the same thing depending on when it is asked and what the result is used for.

Diagnostic assessment

An assessment placed before teaching begins, in order to establish what learners already know. It does not count towards progression, and serves instead to form groups, to adjust the pace of a course, or to identify missing prerequisites before they block the rest of the programme.

Placement tests

Formative assessment

An assessment carried out during learning, whose function is to inform learner and teacher about what is currently being acquired. Its useful output is the feedback, not the mark: it is frequent, low-stakes, and serves to correct course while there is still time to do so.

Formative and continuous assessment

Summative assessment

An assessment that comes at the end of a sequence to take stock of what has been learned and produce a mark that counts. It fixes a result rather than accompanying progress, and differs from formative assessment in its function, not in the format of the questions asked.

Certification assessment

An assessment whose result determines the award of a title, a degree or a certification. The stakes impose stronger requirements on candidate identification, on the fairness of sitting conditions, on the traceability of marking and on how long completed papers are retained afterwards.

Certification and degree exams

Continuous assessment

A mode of assessment that spreads marking across several exercises over time rather than concentrating it in a single final exam. It smooths the effect of one poor performance and gives regular measurement points, at the cost of a heavier organisational and marking load.

Ipsative assessment

An assessment that compares a learner with their own earlier results rather than with a standard or with their peers. It makes individual progress visible, which makes it useful in vocational training and in remedial work, but it allows no ranking between candidates.

Peer assessment

An arrangement in which learners assess one another's work against a supplied rubric. The exercise works on the criteria as much as on the content; its reliability depends closely on how precise the rubric is and on having enough assessors per piece of work.

Self-assessment

A learner's appraisal of their own work against explicit criteria. It develops the ability to judge where one stands, a condition of genuine autonomy, but its results correlate weakly with external assessment for as long as the criteria have not been worked on beforehand.

Testing effect

The phenomenon whereby testing yourself on a piece of content improves your retention of it more durably than rereading it does. Active recall consolidates the memory trace, which makes frequent assessment an instrument of learning in its own right, and not only an instrument of measurement.

Bloom's taxonomy

A classification of learning objectives by level of mental operation, from recalling knowledge through to creating, by way of understanding, applying, analysing and evaluating. It is used to check that an exam does not question only the first of those levels.

Constructive alignment

Coherence between the learning objectives announced, the activities offered and the assessments that close a course. Misalignment is easy to spot: the course works on case analysis, the exam asks for definitions, and so students revise the definitions rather than the analysis.

Learning outcomes

A statement of what a learner must be able to do at the end of a course, written with an observable action verb. Written that way, learning outcomes translate directly into assessment criteria, which an objective phrased in terms of "knowing" does not allow.

Competency-based approach

The organisation of a course and its assessment around competencies — capacities to act in a given situation — rather than around subject content. Assessment then takes the form of complex, contextualised tasks, judged against a rubric rather than on a total of correct answers.

Integrity, proctoring and exam security

What guarantees that the mark really belongs to the candidate — and what each arrangement costs, in money as in personal data.

Remote proctoring

The set of arrangements that make it possible to supervise a candidate sitting outside an exam centre: video and audio capture, screen monitoring, detection of atypical behaviour. The level chosen must stay proportionate to the stakes of the exam and to the data it leads you to collect.

Proctoring at Evalmee

Automated proctoring

Supervision carried out by algorithms that analyse the video feed, the sound and the activity on the device in order to flag atypical events: face absent, second person present, window switched. The flags produced are leads to be investigated, never proof of cheating in themselves.

Live proctoring

Supervision by a human invigilator who follows candidates in real time during the session and can intervene, ask for a sweep of the room or stop the exam. It is the most expensive mode, and the only one that allows an incident to be dealt with as it happens.

Record-and-review proctoring

Recording of the session, reviewed afterwards by an examiner, usually guided by the moments that automated analysis has flagged. It costs less than live supervision and leaves a verifiable trace, but it allows no intervention at all while the exam is being sat.

Browser lockdown

Restriction of the candidate's software environment during the exam: full screen imposed, tab switching blocked or flagged, copy-paste and printing disabled. The measure raises the cost of opportunistic cheating, but can do nothing about a second device sitting next to the computer.

Secure exam browser

A dedicated application that replaces the ordinary browser during the exam and locks down the workstation. It offers stronger control than a simple in-browser restriction, but requires an installation, rules out certain devices and shifts part of the technical support burden onto the institution.

Randomisation

Random drawing of questions from a bank and of the order of the options, so that two neighbouring candidates do not see the same paper. It is the cheapest defence against copying between candidates; it assumes questions of comparable difficulty if it is to stay fair.

Audit trail

A timestamped log of every event in a session: sign-in, display of each question, successive answers, network incidents, invigilator actions. It is what makes it possible to reconstruct a contested exam and to answer a complaint with something other than memory.

Candidate authentication

Verification that the person sitting the exam is the person enrolled: named credentials, sign-in through the institution's directory, checking an identity document, sometimes a biometric comparison. The level of verification is set by the stakes of the exam and by the data it makes it legitimate to process.

Exam misconduct

Any conduct aimed at obtaining a mark that does not reflect the candidate's own work: unauthorised documents, communication with a third party, impersonation, undeclared use of a generative tool. How it is classified and what follows from it are set by the institution's regulations.

Plagiarism

Taking someone else's work without acknowledging the source, and presenting it as your own production. Similarity tools measure overlaps of text, which is a lead and not a verdict: classifying something as plagiarism requires an examination of intent and of context.

Interoperability and technical standards

The standards that let an assessment platform plug into an institution's information system rather than live alongside it.

LMS (learning management system)

A platform that hosts an institution's courses, enrolments and learner tracking. It is students' daily point of entry, which makes integration with the LMS — rather than the standalone quality of an exam tool — the practical condition of that tool actually being adopted.

Moodle

An open-source learning platform, distributed under the GPL licence and very widely deployed in French-speaking higher education. Its openness explains why it serves as the reference for integration: an assessment tool that does not plug into it forces teachers to keep two lists of students.

Source: moodle.org

LTI Advantage

A set of three services added to LTI 1.3 by 1EdTech: Assignment and Grade Services, which creates the gradebook columns and posts results back into them; Names and Role Provisioning Services 2.0, which exposes users and their roles; Deep Linking 2.0, which selects specific content inside the external tool.

Source: 1EdTech — LTI Advantage

SCORM

A set of technical specifications published by the ADL Initiative to make e-learning content interoperable and reusable from one platform to another. It describes how content is packaged and the run-time interface through which a module reports learner progress to the LMS hosting it.

xAPI (Experience API)

IEEE standard 9274.1.1-2023, derived from the ADL Initiative's xAPI specification, which defines a JSON data model and a web interface for recording learning experiences to a Learning Record Store. Unlike SCORM, it also records activities that took place outside any LMS.

Source: IEEE 9274.1.1-2023

QTI (Question and Test Interoperability)

A 1EdTech standard describing an exchange format for questions, tests and results. It allows a question bank to be carried from one system to another without re-keying it, which makes a library of items independent of the platform that produced it.

Source: 1EdTech — QTI

Single sign-on (SSO)

A mechanism allowing a user to sign in to several services with their institutional account, through protocols such as SAML, CAS, Shibboleth or OpenID Connect. It removes the need to create dedicated passwords and makes revoking an access immediate when an account is closed.

SSO and connectors

REST API

A programming interface through which a third-party system reads and writes a platform's data using HTTP requests. It makes it possible to automate what the graphical interface does by hand: creating sessions, enrolling candidates, posting results into a student records system.

Accessibility and inclusion

What French law requires of an institution's digital services, and the adjustments a candidate is entitled to.

Digital accessibility

The design of a digital service so that it can be used by people with disabilities: compatibility with screen readers, full keyboard navigation, sufficient contrast, content that is never conveyed by sight or sound alone. It is verified by audit, not by a statement of intent.

RGAA (French digital accessibility standard)

Référentiel général d'amélioration de l'accessibilité, the reference standard issued under article 47 of the French law of 11 February 2005. It translates the legal obligations of digital accessibility into control criteria backed by technical tests; its version 4 was adopted on 20 September 2019.

Source: accessibilite.numerique.gouv.fr

Accessibility statement

The public document that a digital service subject to the RGAA must publish to declare its state of compliance. Three states exist: full compliance when every criterion is met, partial compliance from half the criteria upwards, and non-compliance below that or in the absence of a valid audit.

Source: accessibilite.numerique.gouv.fr — accessibility statement

Tiers-temps (French statutory extra exam time)

The extension of exam time granted to a candidate with a disability. The French code de l'éducation provides that it may not exceed a third of the time normally allowed, a longer extension remaining possible exceptionally, on the reasoned request of the doctor who examines the situation.

Source: Code de l'éducation, articles D351-27 to D351-31

Aménagements d'épreuves (French exam access arrangements)

The full set of adaptations available to a candidate with a disability: adapted material conditions, technical and human assistance, spreading exams over several sessions, retention of marks for five years, adapted or waived papers. The authority organising the exam decides and notifies its decision.

Source: Code de l'éducation, articles D351-27 to D351-31

Universal design for learning

An approach that plans, from the design stage, several ways of presenting content, of engaging with a task and of demonstrating what has been learned, rather than adding adjustments after the fact. Applied to assessment, it reduces the number of situations calling for individual treatment.

French regulatory and quality framework

The bodies and registers that give a French certification its value, and the data protection obligations that bear on any online exam.

Qualiopi (French training quality certification)

A quality certification mark issued by an accredited certification body, covering the processes of providers delivering activities that contribute to skills development. Since 1 January 2022 it conditions access to public or pooled funding, and it rests on the Référentiel national qualité, the national quality framework.

Source: Ministère du Travail — Qualiopi

RNCP (French national register of professional certifications)

The Répertoire national des certifications professionnelles, drawn up and kept up to date by France compétences. The certifications registered in it are classified by qualification level and by field of activity, and they are made up of competency units that can be assessed and validated separately.

Source: France compétences — professional certification

Répertoire spécifique (French register of complementary certifications)

A second national register drawn up by France compétences, which lists the certifications and authorisations corresponding to professional skills that complement professional certifications. Registration follows an application, after the opinion of the competent commission, for a maximum period of five years.

Source: Code du travail, article L6113-6

France compétences (French vocational training authority)

A public administrative body created on 1 January 2019 by the French law of 5 September 2018, and the sole national governance authority for vocational training and apprenticeship. Its governance is quadripartite and its remit is to fund, to regulate and to improve the training on offer.

Source: France compétences — its role

Bloc de compétences (French competency unit)

Under the French code du travail, a homogeneous and coherent set of competencies contributing to the autonomous exercise of an occupational activity, and capable of being assessed and validated separately. Each certification is defined by an activity framework, a competency framework and an assessment framework.

Source: Code du travail, article L6113-1

Modalités de contrôle des connaissances (French assessment regulations)

The assessment rules adopted by a French higher education institution. The code de l'éducation provides that they be set no later than the end of the first month of the teaching year and may no longer be changed during that year, assessment being by continuous assessment, by a final exam, or by both.

Source: Code de l'éducation, article L613-1

Hcéres (French higher education and research evaluation authority)

Haut Conseil de l'évaluation de la recherche et de l'enseignement supérieur, an independent public authority in France. It evaluates higher education institutions, research organisations and research units, as well as the programmes they run and the degrees they award, following an evaluation methodology based on peer review.

Source: hceres.fr

CTI (French engineering degree accreditation commission)

Commission des titres d'ingénieur, an independent body charged by French law since 1934 with evaluating engineering schools with a view to accrediting them, with developing the quality of their programmes and with promoting the title of engineer. It issues an opinion for public schools and a decision for private ones.

Source: cti-commission.fr

CEFDG (French management degree evaluation commission)

Commission d'évaluation des formations et diplômes de gestion, the sole national body competent to evaluate the quality of programmes at private and chamber-of-commerce management schools. Under joint ministerial supervision, it examines state-endorsed degrees and those that confer the grade of licence or master.

Source: cefdg.fr

Data minimisation

The GDPR principle under which, in the CNIL's wording, personal data must be adequate, relevant and limited to what is necessary in relation to the purposes for which it is processed. Applied to remote proctoring, it requires collecting only what the stakes of the exam justify.

Source: CNIL — data minimisation

Analyse d'impact (AIPD, French data protection impact assessment)

Analyse d'impact relative à la protection des données: according to the CNIL, a tool for building a processing operation that complies with the GDPR and respects privacy. It is required for processing operations likely to result in a high risk to the rights and freedoms of the people concerned.

Source: CNIL — AIPD

Data retention period

The length of time personal data may remain in an active database. The CNIL points out that it is determined by the purpose of the processing: at the end of that period the data must be deleted, anonymised or archived, which bears directly on remote proctoring recordings.

Source: CNIL — retention periods

Frequently asked questions about assessment vocabulary

What is the difference between formative and summative assessment?

The difference lies in the function, not in the format. A formative assessment happens during learning and serves to inform the learner and the teacher about what is currently being acquired: its useful output is the feedback. A summative assessment happens at the end of a sequence to take stock and produce a mark that counts. The same MCQ can serve both.

What does the discrimination index of a question measure?

It measures a question's ability to tell strong candidates from weak ones, by comparing the success rate on that item among those with the highest overall scores and among those with the lowest. An index of zero flags a question that carries no information; a negative index almost always flags a wrong answer key or an ambiguous wording.

Does an exam platform have to comply with the RGAA?

The RGAA applies to the digital services of public bodies and of certain companies, under article 47 of the French law of 11 February 2005. An institution that runs its exams on a platform remains responsible for the accessibility of the service it offers its students: the tool's compliance is therefore a purchasing criterion, and it is checked against the accessibility statement published by the vendor.

What does LTI 1.3 add over a simple link to an LMS?

LTI 1.3 is a 1EdTech standard: the external tool is launched from the LMS with no separate account, and the security of the exchange rests on OAuth 2.0 and JSON Web Token credentials. With the LTI Advantage services, the LMS also passes on the list of enrolled users and their roles, and pulls marks back into its gradebook automatically — which a plain hyperlink does not do.

Is extra time limited to a third of the exam duration?

The French code de l'éducation provides that the extension of exam time may not exceed a third of the time normally allowed. A longer extension remains possible exceptionally, on the reasoned request of the doctor who examines the candidate's situation. It is the administrative authority organising the exam that decides which adjustments are granted and notifies its decision to the candidate.

Put this vocabulary to work

Criterion-referenced marking schemes, item analysis, extra time: Evalmee gives you the tools for what this glossary describes.