Typical mistakes in Horizon Europe proposals, and what evaluators actually write
Horizon Europe has become a near-zero-margin game. As the number of applications has surged, the bar has risen with it: on a competitive RIA or IA, a score that would have been funded a few years ago now misses the cut. You are no longer writing to pass. You are writing for something close to 15 out of 15, and one weak section is enough to end it.
And proposals rarely fail on a single big mistake. They fail on a handful of specific, recognisable weaknesses that come back, evaluation after evaluation, in almost the same words. There are two kinds, and telling them apart is the whole game.
Erosion is the first: small shortcomings, each shaving a fraction of a point. A few are survivable, but do not count on it. The evaluators' own rulebook is explicit that they compound: a proposal "with a large number of shortcomings, all together" can itself amount to a significant weakness and fall below the threshold. Erosion is not safe, it is just slower.
A significant weakness is the second, and the faster route down: one structural failure that, on its own, addresses a criterion "in a limited and/or not sufficiently effective way", dropping it below the threshold of 3 and taking the proposal with it, however strong the other thirty-nine pages are. In one real evaluation, Excellence came out at 2.5, below threshold, from a single gap: the methodology was thorough and well written, but never included the one measurement the call's science required. Everything else was beside the point.
This page walks the weaknesses that recur, criterion by criterion and axis by axis, following the same axes a Horizon Europe panel works through. The examples are drawn from real evaluations and lightly anonymised: the project name, the partners and the technical domain have been removed, and only the pattern is kept. For each: what the proposal wrote, what the evaluator saw, and what to do instead.
Excellence
Methodology completeness: the axis that broke the floor
Significant weakness. Below threshold on its own.
What the proposal wrote (anonymised):
"The methodological approach can be visualised as a continuous cycle of four main blocks. (i) Data acquisition gathers historical records, in-line monitoring at pilot sites, sensor-stream data, and external open datasets. (ii) Modelling uses machine-learning algorithms and state-of-the-art domain models. (iii) Assessment and operator uptake, where results are integrated by operators during the project. (iv) A decision-support system and recommendations, delivered through an interactive platform and a mobile application. The key benefit is that operator uptake occurs within the project, so the methodological loop runs at least once."
Four clean blocks, each described. It reads complete.
What the evaluator saw: "The methodology is descriptively complete but omits critical elements required for the domain and the call scope." The proposal never specified how the one measurement at the heart of the call's science would be produced and fed into the models. The central concept was "asserted but not operationalised", and the proposal "conflates monitoring (data collection) with the research activity itself, failing to demonstrate how the project will advance understanding beyond observational data collection."
Why it matters: a methodology can be complete as a description and empty as a method. This one was well written and lost nothing on style, and it still scored 2.5, below threshold, on this axis alone. If the call's science requires a specific analysis, it must appear by name, with how its output feeds your models. Describing your data pipeline is not describing your method.
State of the art and novelty
Minor shortcoming. A point-shaver, unless novelty is the heart of your case.
What the proposal wrote (anonymised):
"The project aims to create a comprehensive decision-support system by strategically integrating multiple cutting-edge technologies, advancing the technology from TRL 4 to a validated TRL 6, creating a far more powerful solution than its individual components."
What the evaluator saw: the claimed breakthroughs "remain largely conceptual. The advancement over the state of the art is asserted rather than demonstrated with concrete technical differentiators," with no measurable gap against the closest existing systems, which the evaluator then named.
What to do instead: "Strategically integrating cutting-edge technologies" describes ambition, not novelty. Name the closest existing systems, including the commercial ones, and quantify the gap you close.
KPI quantification
Minor shortcoming. Point-shaver, but it recurs on every criterion.
What the proposal wrote (anonymised): most objectives carried real numbers, for example a prediction-error target expressed as a percentage reduction against a stated baseline and a response-time target. But one objective read only "validation of a coefficient", and the societal outcomes were phrased as "improved efficiency" and "better" conditions.
What the evaluator saw: "Most objectives are quantified with specific target values. However, [the weak one] lacks a clear definition or baseline value, making it difficult to assess whether the target is ambitious or merely descriptive," and the societal outcomes had no time-bound indicators.
What to do instead: The evaluator lands on the weakest objective, not the average. One unquantified objective among six good ones is exactly where the comment goes, and a target with no baseline cannot be judged ambitious. Give a baseline, a target and a deadline to every objective, especially the societal ones.
Call scope coverage: a structural axis
Minor shortcoming when a few activities are thin, but a significant weakness the moment a required activity is left uncovered. Structural axes swing between the two.
A proposal typically opens its alignment section by stating that it "directly contributes" to the call's objectives, then describes its own activities. The evaluator's response is the recurring one: the proposal "broadly aligns with the call scope but fails to address [specific required activities] explicitly", then names the activities the call required and the proposal never covered.
"Directly contributes" is an assertion, and assertions are what evaluators strike out. The call lists its activities; map each one to a work package or objective, in a table, and cover them all. This axis is structural: failing it does not shave a point, it can break the floor.
Two more Excellence axes catch teams out: the gender dimension (acknowledged in a paragraph, never used as an analytical variable in the research design) and open science and data management (mandatory practices committed to, but no plan for the sensitive or proprietary data you will actually produce).
Impact
Barriers and mitigation
Minor shortcoming here (risks were named), but a significant weakness when the barriers section is empty or absent.
What the proposal wrote (anonymised), in its risk table:
"Stakeholder uptake risk (low adoption of the system). Level: High. Probability: High. Mitigation: engage users from the start, co-design features, frequent demonstrations, feedback loops, ensure simplicity and relevance to end-users."
What the evaluator saw: the proposal "identifies technical and adoption risks and provides plausible mitigation directions. However, it does not systematically address dominant structural barriers," which the evaluator listed: regulatory fragmentation across member states, missing infrastructure, user scepticism toward automated recommendations, and established commercial competitors.
What to do instead: Do not confuse project risks with barriers to impact. Risks sit inside your project. Barriers sit outside it: regulation, market, behaviour, competition. Co-design is not a mitigation for regulatory fragmentation.
Exploitation and key exploitable results
Minor shortcoming. A point-shaver, and a common one for an IA where exploitation weighs heavily.
What the proposal wrote (anonymised):
"The consortium is committed to a clear and innovative path toward both open access and commercialising its innovations, ensuring long-term sustainability while maintaining a commitment to public benefit."
What the evaluator saw: the key exploitable results were listed and partners named, but "the exploitation plan is generic and lacks specificity," with routes to uptake (licensing, spin-off) not detailed and IP management described in broad terms.
What to do instead: "Committed to a clear path" is the tell. Each result needs a named owner, a TRL target, a concrete route to uptake and an IP position, and the plan must address each target group's different needs rather than the consortium's own.
KPI quantification, again: where Impact quietly loses
Minor shortcoming, on the criterion where it costs the most.
The same pattern as in Excellence, in the criterion where it costs most: teams give precise numbers for the technical outputs they are comfortable with, and go qualitative for the societal outcomes, which is exactly what the Impact criterion scores. Every outcome needs a baseline, a target and a deadline, especially the soft ones.
Expected outcomes coverage: the structural axis, and the single most overlooked failure
Significant weakness when the pathway is not mapped to the call's outcomes. This is the classic below-threshold Impact score.
The counterpart of call scope coverage, on the Impact side. The call's expected outcomes get referenced in passing, usually inside the consortium description, but are never mapped result by result to what the project delivers. The evaluator's phrase is always some version of asserted rather than demonstrated. Because this axis is structural, an unmapped pathway does not cost a fraction of a point: it takes Impact below threshold. Quote the expected outcomes verbatim and answer them one at a time.
Quality and efficiency of the implementation
Consortium expertise: the same hole, a second time
Significant weakness. It undermines the central innovation claim, and does so on a second criterion.
What the proposal wrote (anonymised):
"The project is powered by a robust, multi-country consortium. This diverse group synergises expertise across three key pillars: advanced technology development, industrial deployment, and techno-economic analysis."
What the evaluator saw: "The consortium lacks demonstrated expertise in [the one sub-discipline the project's central claim depended on]. No partner is named as having prior work in [that specific method], despite this being a core innovation claim."
Why it matters: this is the same missing competence as the methodology gap in Excellence, and it is not a coincidence. A hole in a load-bearing capability weakens two genuinely different things: whether the method is sound (Excellence) and whether this consortium can deliver it (Implementation). Evaluators are instructed not to mark the same point down twice, and they do not have to: a core competence that nobody in the consortium owns is a real, separate weakness on each criterion. Listing disciplines is not demonstrating the one your central claim rests on. Name it, and give it a partner with a track record.
Resource and effort allocation: the table is a statement of priorities
Minor shortcoming, unless the mismatch with the ambition is gross.
What the proposal wrote (anonymised): an effort table spreading several hundred person-months across nine work packages, with validation and piloting each drawing materially more effort than the work package building the predictive models at the analytical core of the project.
What the evaluator saw: allocation "globally proportional", but the core work package "receives less than field validation or piloting. This under-resourcing is concerning given its technical ambition," and no rationale was given.
What to do instead: The panel reads your effort table as a statement of priorities. If the work package carrying your novelty is resourced below the peripheral ones, the table contradicts the ambition in section 1, and the panel believes the table. Justify the split.
The remaining implementation axes are where the small shortcomings accumulate: risk mitigation (generic risks with generic mitigations), work package timing and dependencies, work plan coherence, procedural compliance (a data management plan scheduled after data collection starts, too few milestones to monitor a long project), and management and governance (a coordinator and a steering committee, with no decision-escalation mechanism sized for the consortium).
The habit that prevents most of this
Read your own draft in the evaluator's voice, axis by axis, and be honest about which sentences are assertions. Almost every comment on this page attaches to one: "directly contributes", "committed to a clear path", "strategically integrating", "the loop runs at least once". Each describes an intention where the panel needed evidence.
Then separate the two failures. The small ones cost fractions of a point. The structural ones, call scope coverage, expected outcomes coverage, and a methodology missing what the call's science requires, take you below threshold on their own. Fix those first.
That is what GrantForge's pre-submission evaluation does: it scores your proposal on these exact axes the way real evaluators do, quotes the passage that will draw the comment, and separates the point-shavers from the threshold-breakers, while you can still act on them.
Run a pre-submission evaluation in GrantForge →
Part of our guide on how to write a winning Horizon Europe proposal. See also how Horizon Europe proposals are scored.