How Horizon Europe proposals are scored, and how funding is really decided

Every Horizon Europe evaluation ends in three numbers. Not a narrative verdict, not a general impression: three scores from 0 to 5, one per criterion, produced by a rubric the European Commission publishes and almost nobody studies. Consortia spend months polishing prose, and minutes, if that, on the scale that decides them.

This page reproduces the real rubric: the official 0 to 5 definitions, the thresholds, the Impact weighting for Innovation Actions, and the part almost no web page carries, the official tie-break order that decides which of two equally scored proposals gets funded. Scope: collaborative Horizon Europe calls under Pillar 2, Research and Innovation Actions (RIA) and Innovation Actions (IA). Success rates are low, which means most funding decisions happen at the margin, exactly where these mechanics bite hardest.

The three criteria, exactly as evaluators score them

RIA and IA proposals are scored against the same three criteria, and the official wording is more precise than most summaries suggest.

Excellence asks for the "Clarity and pertinence of the project's objectives, and the extent to which the proposed work is ambitious and goes beyond the state of the art", together with the soundness of the methodology: concepts, models, assumptions, interdisciplinarity, the gender dimension in the R&I content, and open science practices.

Impact asks two questions: the "Credibility of the pathways to achieve the expected outcomes and impacts specified in the work programme", and the "Suitability and quality of the measures to maximise expected outcomes and impacts, as set out in the dissemination and exploitation plan, including communication activities."

Quality and efficiency of the implementation covers the "Quality and effectiveness of the work plan, assessment of risks, and appropriateness of the effort assigned to work packages, and the resources overall", plus the "Capacity and role of each participant, and the extent to which the consortium as a whole brings together the necessary expertise."

Each criterion gets its own score and its own threshold. A brilliant Excellence section cannot rescue a weak Impact section: the scores do not compensate each other at threshold level.

The 0 to 5 scale, verbatim

Each criterion is scored from 0 to 5, in steps of 0.5. The official definitions:

Score Label Official definition
0 Fails to address the criterion or cannot be assessed.
1 Poor Inadequately addressed, or serious inherent weaknesses.
2 Fair Broadly addresses the criterion, but there are significant weaknesses.
3 Good Addresses the criterion well, but a number of shortcomings are present.
4 Very good Addresses the criterion very well, but a small number of shortcomings are present.
5 Excellent Successfully addresses all relevant aspects; any shortcomings are minor.
Each criterion is scored 0 to 5, in steps of 0.5 1 2 3 4 5 Poor Fair Good Very good Excellent threshold 3 below threshold, criterion fails at or above threshold 3 per criterion · 10 overall (three 3s sum to 9, which fails)
The scale evaluators score on, and the threshold that decides you.

Read the ladder carefully: the difference between 3 and 4 is not quality of writing, it is the count and weight of shortcomings. Evaluators are not asking "is this good work?", they are asking "how many things are wrong, and how badly wrong are they?"

The thresholds: 3 per criterion, 10 overall

Verbatim from the official rules: "The threshold for the individual criteria will be 3. The overall threshold, applying to the sum of the three individual scores, will be 10."

The arithmetic has a trap in it. Three scores of exactly 3 sum to 9, which fails the overall threshold of 10. Merely "Good" everywhere is a rejection. You need at least one criterion above 3, and in practice, given how few proposals get funded, you need well above threshold on all three.

And passing thresholds is not funding. It only puts you on the ranked list, where the call budget runs out long before the list does.

Shortcoming or significant weakness: the line that decides your score

The single most useful distinction in the whole rubric is the one between a shortcoming and a significant weakness.

A significant weakness, in the official briefing given to every expert, "means the proposal addresses the criterion in a limited and/or not sufficiently effective way (will lower the score below threshold). This can also be the case when the proposal includes a large number of shortcomings ... all together."

Two consequences. First, one significant weakness is enough to sink a criterion below 3, whatever else the section does well. Second, the accumulation rule: a pile of individually small shortcomings can be treated, together, as a significant weakness. Death by a dozen cuts is an official scoring outcome, not a metaphor.

The briefing is blunt about the ceiling too: "Proposals with significant weaknesses that prevent the project from achieving its objectives ... must not receive above-threshold scores."

Here is what that looks like on a real, anonymised proposal.

What the proposal wrote (methodology section): "The methodological approach can be visualised as a continuous cycle of four blocks: data acquisition, modelling, assessment and stakeholder uptake, and a decision-support system. The key benefit is that stakeholder uptake occurs within the project, so the methodological loop runs at least once."

What the evaluator saw: "The methodology is descriptively complete but omits critical elements required for the domain and the call scope." The proposal never specified how the one measurement the call's science required would be produced and fed into the models. The central concept was "asserted but not operationalised".

The score it produced: Excellence 2.5, below the threshold of 3. Not from an accumulation of small issues: from one structural gap in a section that was otherwise well written. That is a significant weakness, and it is why "well written" is not the same as "scores well".

How evaluators actually read

The official briefing tells experts to evaluate each proposal "as submitted and not on its potential if certain changes were to be made". There is no benefit of the doubt: what you meant, what your team obviously knows, what "any expert would infer", none of it counts. Only what is on the page.

Experts also "explain shortcomings, but do not make recommendations". The ESR you receive will tell you what was wrong, never what to do about it.

In practice, evaluators navigate your proposal with the evaluation form as their map: they hunt for the explicit answer to each sub-criterion, criterion by criterion. One small mercy is built in: they are instructed not to mark the same critical aspect down twice under two different criteria. But a point the form asks for and your text never explicitly answers is, for scoring purposes, absent.

The process itself is short to describe: several independent experts each read and score your proposal (each writing an individual evaluation report), then they meet to reach a consensus, which is written up as the Evaluation Summary Report you receive. Many of the typical mistakes in Horizon Europe proposals are, at root, failures to write for this reading mode: text organised around the applicant's story instead of the evaluator's form.

The Impact weighting, and what changed in 2026

For Innovation Actions, one more rule shapes the outcome. Verbatim: "To determine the ranking for Innovation actions, the score for Impact will be given a weight of 1.5."

Two nuances that most guides blur. The weighting applies to Innovation Actions only, not to RIAs. And it applies only to the ranking: it is never used to decide whether you passed the thresholds. An IA still needs a raw 3 on Impact; but once on the ranked list, its Impact score counts one and a half times. For an IA, the Impact section is where ranking positions are won and lost.

One currency note, because most web guidance has not caught up. From Work Programme 2026/2027 onwards, the evaluation of Impact no longer considers the "scale and significance" of the contributions. Impact is now judged on the credibility of the pathway and the suitability of the dissemination, exploitation and communication measures. Inflating projected numbers buys nothing; a believable causal chain is the whole game. (A minor related change: Do No Significant Harm is now required only for EIC Accelerator topics.)

The tie-break order: how funding is decided at the margin

Call budgets are fixed, so somewhere on every ranked list there is a line: above it, funded; below it, not. Around that line sit proposals with identical scores, and the official General Annexes define exactly how ties are broken, in sequence:

  1. Proposals that cover aspects of the call not covered by higher-ranked proposals.
  2. Then the score for Excellence (for Innovation Actions: the score for Impact first, then Excellence).
  3. Then gender balance among the researchers primarily involved in the proposal (the core research team).
  4. Then geographical diversity.
  5. Then portfolio synergies and SME participation.

Point 1 is the actionable one, and almost nobody plans for it. If two proposals score the same, the one that covers an aspect of the call that no higher-ranked proposal covers wins the tie. Which means that deliberately owning a required but less-crowded aspect of the topic, and covering it visibly, is a real positioning move, not decoration. The crowded centre of a topic is where ties happen; the neglected edge of its scope is where they are broken.

Two-stage calls play by different numbers

If your call is two-stage, stage 1 uses a different bar. Only Excellence and Impact are evaluated, the threshold for each is 4, not 3, and the overall threshold is normally 8 or 8.5. A "Good" section that would survive a single-stage evaluation fails at stage 1. Short first-stage proposals are not a lighter exercise: they are a harder one, judged on two criteria at "Very good" level.

Score yourself before they do

The rubric is public, the definitions are verbatim, and the tie-break order is written down. Nothing about Horizon Europe scoring is secret; it is just rarely applied to a draft before submission, when it can still change the outcome.

Read your own Part B the way the experts will: as submitted, form in hand, hunting for the explicit answer to each sub-criterion, counting shortcomings, and watching for the one structural gap that turns a well-written section into a 2.5.

→ Run your proposal through GrantForge's pre-submission evaluation: a 0 to 5 score per criterion in the evaluators' own language, with the threshold-breaking weaknesses separated from the fraction-of-a-point shortcomings, while there is still time to fix them.


Part of our guide on how to write a winning Horizon Europe proposal. See also typical mistakes in Horizon Europe proposals.