Hollow Knight: Silksong · boss difficulty survey

The Needle'sVerdict

Sixteen thousand players scored every boss in Silksong. The ranking they produced is one of the most statistically solid things a community poll has ever produced — and the interesting part is not who came first.

Responses
Ratings cast
Bosses rated
Avg. completion
Reliability

How to read the numbers

Found it hard
The share of players who scored a boss 4 or 5 out of 5. This is the main number used throughout — it says what fraction of people struggled, which is easier to picture than an average.
Found it easy
The share who scored it 1 or 2.
Divided score
0 to 100. How evenly the vote spread across the five options. Near 0 means near-unanimous; near 100 means the community produced no answer at all.
Missed it
The share who answered "Did not fight" — a rough measure of how easy a boss is to walk past.

96.9% found the hardest boss hard

A ranking you can actually trust

Community polls usually produce noise dressed up as a result. This one didn't. 16,361 people rated up to 46 bosses each, and the margin of error on any single boss's score lands between 0.004 and 0.014 points — far smaller than the gaps between them.

Margin of errorHow far a score might shift if you polled a different 16,361 people. Below about 0.05 here, which is why the order holds.

So the ordering below is closer to a measurement than an opinion. Read it by the bars rather than the ranks: the bar shows how the whole vote split, and that shape carries more than any average can.

Fig. 1 — Every boss, by how 16,361 players scored it

Scored 
1 — trivial  →  5 — brutal
Act 1
Act 2
Act 3
← found it easyfound it hard →
Hard
%
Avg
Bars are centred on the middle of the scale: everyone who answered 1 or 2 extends left, everyone who answered 4 or 5 extends right, and the 3s straddle the line. A bar leaning right was a hard fight; leaning left, a walk-through. Hard % is the share scoring 4 or 5. Hover any row for the full breakdown.

Two independent halves agree at 0.9999

How much of this is signal

Worth testing before drawing conclusions from it. Split the 16,361 respondents at random into two halves of roughly 8,180 and let each half rank the 46 bosses independently. Do that 200 times. The two rankings agree at a correlation of 0.9999 every single time, and the worst run still lands at 0.9998.

What 0.9999 meansTwo completely separate crowds of 8,180 players produced, for practical purposes, the same list. The result is not an artefact of who happened to answer.

That level of agreement means the survey was enormously oversampled for the question it asked. Working backwards: how many respondents would you actually have needed?

Fig. 2 — How many players it takes to get this ranking

Rank agreement between a random subsample and the full 16,361-person result. 250 people already reproduce the ranking at 0.97; past about 1,000 the curve is flat. The extra 15,000 responses buy confidence in the close calls, not the overall shape.

Resampling the data 1,000 times puts a confidence range on each boss's position. 35 of the 46 never move at all — their rank is identical in every resample. The remaining 11 shift by one or two places, and always with an immediate neighbour.

Widest uncertaintyPinstress, Bell Eater and Raging Conchfly share ranks 24–26 and cannot be separated. Nothing in the data moves further than two positions.

The honest version of the ranking is therefore slightly shorter than 46. Testing each neighbouring pair on just the players who fought both, 15 pairs fail to separate. Collapsing those leaves 31 genuinely distinguishable positions, with twelve tied blocks:

Fig. 3 — Bosses that cannot be separated from each other

Adjacent pairs where a paired test on players who fought both bosses returns p > 0.001. Anyone arguing about the order inside one of these blocks is arguing about noise.

87.4% gave First Sinner the maximum

The scale ran out before the game did

A five-point scale can only measure what fits inside it, and seven bosses broke out of the top. First Sinner was scored 5 out of 5 by 87.4% of everyone who fought it. Skarrsinger Karmelita by 84.6%, Lost Lace by 76.2%, Cogwork Dancers by 71.1%. Counting everyone who said 4 or 5, First Sinner reaches 96.9% — essentially the whole playerbase.

This is a ceiling effect, and it changes how the rest of the table should be read. The gap between First Sinner (4.82) and Widow (4.53) looks trivial next to the gap between Widow and Bell Beast (3.19). It almost certainly isn't. Once four-fifths of your sample has chosen the maximum, the instrument has stopped measuring — the real distance between Silksong's hardest fights is compressed into three tenths of a point because there was nowhere else for those players to put their answer.

The top of this ranking is a floor, not a measurement. We know First Sinner is at least as hard as the scale can express. We do not know how much harder.

n = 16,156 players who fought First Sinner

The same wall exists at the bottom, less dramatically: 85.0% found Plasmified Zango easy, and 77.9% said the same of Palestag.

Boss identity 49.4% of variance · rater identity 7.8%

The community agrees far more than it argues

Break down all 705,401 ratings and the split is clear. Differences between bosses account for 49.4% of the variation. Differences between people — the harsh-grader effect — account for 7.8%. Which boss you are scoring matters about six times more than who is doing the scoring.

ConsistencyCronbach's α across the 46 items is 0.885. Above 0.8 is normally considered a coherent measuring instrument rather than a collection of unrelated opinions.

Re-scoring every rating relative to that respondent's own average — which strips the harsh-grader effect out entirely — changes the rank order of exactly zero bosses. But agreement is not uniform, and where it breaks down is the interesting part.

Fig. 4 — Consensus collapses in the middle of the difficulty range

Act 1
Act 2
Act 3
The vertical axis is the divided score: 100 would mean the vote split evenly across all five options, 0 that everyone picked the same number. Bosses at both extremes of difficulty are near-unanimous. The genuine arguments live in the middle of the range, and one boss sits at the very top of the arch.

Groal the Great · divided score 94 of 100

Groal the Great is the most contested fight in Pharloom

Groal's votes are almost perfectly flat: 26.0% said 1, 22.1% said 2, 28.5% said 3, 16.9% said 4, 6.5% said 5. Its divided score of 94 is close to the theoretical maximum. Sixteen thousand people fought this boss and collectively produced no answer.

Flat, not merely wideA boss can average 3 because everyone agrees it is average, or because half the players say 1 and half say 5. Only the second is disagreement, and only Groal really looks like it.

That flatness is not what "medium difficulty" looks like. A genuinely medium boss looks like Bell Beast, where 54.1% of players chose exactly 3 and the divided score is the lowest in the dataset outside the extremes. Groal isn't medium. Groal is two different fights depending on who you ask — and it turns out you can find the two groups.

5,978 vs 3,924 players · 1.2-point disagreement

The playerbase splits into two camps, over three fights

Ignore how hard each player found the game overall, and group people purely by the shape of their answers — which bosses they found relatively harder than the rest. Restricting this to the 9,902 respondents who fought at least 44 of the 46 bosses, so that missing answers can't drive the result, the playerbase falls into two clean groups of 5,978 and 3,924.

The two camps disagree about almost nothing. Across 43 of the 46 bosses their scores sit within about half a point. Then there are three fights.

Fig. 5 — Where the two camps disagree

Camp A — 5,978 players
Camp B — 3,924 players
Average score from each camp, for the 18 bosses where they differ most. Savage Beastfly, Groal the Great and Broodmother separate the two groups by up to 1.3 points — several times any other fight. Below those, the pattern is a mild general tendency: Camp A uses more of the scale, Camp B compresses toward the middle.

That general tendency is real but modest — it accounts for 55% of the difference between the camps, and once you subtract it, the three fights at the top are still five times further out than any other boss. They are not an artefact of some people being more generous with 1s.

How the gap was checkedCamp difference regressed on how far each boss sits from the midpoint. Residuals: Savage Beastfly II +0.82, Groal +0.82, Savage Beastfly I +0.79, against a spread of 0.16 for the other 43.

So the single largest genuine disagreement in Silksong's playerbase is about the two Savage Beastfly encounters and Groal the Great — three fights that roughly 60% of players found trivial and 40% found middling. For Savage Beastfly I the two camps average 1.70 and 2.98. For the Far Fields rematch, 1.59 and 2.92.

So what do these three fights have in common? Cross-referencing a published boss guide gives an answer that fits almost too neatly: each one has a documented way to trivialise it. Groal can be fought from under the lower-right platform, where none of his attacks reach. Savage Beastfly can be lined up so that its own vertical charge kills the allies it summons. Both are described in strategy guides as cheese.

Fig. 6 — Bosses with a documented trick, against everything else

How far the two camps' scores differ, for the six bosses a strategy guide gives an explicit trivialising trick versus the other 40. The cheeseable fights split the playerbase by 0.88 points; everything else by 0.04 (p = 0.004). Five of the six biggest splits in the game are cheeseable fights.

That is a much better fit than difficulty itself. A fight you either walk through or grind out produces exactly this shape — two humps, not a spread — because knowing the trick is binary. It also explains Groal's flat vote: the 26% who scored it 1 and the 24% who scored it 4 or 5 were, in effect, playing different fights.

Correcting an earlier readAn earlier draft of this report guessed the divide was about airborne bosses, since Moorwing and the Conchflies appear high in the list. The cheese pattern fits better and is externally documented; the airborne overlap is largely coincidental, as Moorwing also has a documented cheese.

The survey asked nothing about tools, routes or tactics, so this remains an inference from an external source rather than a measured result. But it is a testable one: a follow-up poll that asked how people beat Groal would settle it in one question.

17 of 17 significant gaps point one way

Players who did more of the game found it harder

Split respondents by how much of the game they got through — 45 or more bosses fought versus 40 or fewer — and compare their scores on the 32 bosses that almost nobody skipped, so both groups are rating the same fights.

Fig. 7 — How thorough players and partial players disagree

Difference in average score between players who fought 45+ bosses (n = 6,624) and those who fought 40 or fewer (n = 1,960). Every difference that clears the significance bar is shown — all 17 fall on the same side of zero.

Two things stand out. The left half of that chart is empty: of the 32 bosses eligible, not one was rated significantly easier by the players who had done more. And the effect vanishes entirely on the earliest mandatory fights — Bell Beast, Fourth Chorus, Lace I, Moss Mother I, Cogwork Dancers, Skull Tyrant all sit indistinguishably at zero.

Direction of the surpriseThe intuitive prediction is the opposite: more experience, lower scores. The data says thorough players rate the optional and late-game fights up to half a point harder.

The pattern is that the gap opens exactly where players' circumstances diverge. Where the game funnels everyone through the same door at the same power level, the scores converge. Where it doesn't — the optional, out-of-the-way, late fights — they split. Groal again leads, at half a point.

This is an interpretation of the pattern rather than something the survey measured; it did not ask about routes, tools or build. An alternative reading is simply that thorough players remember more of the struggle, having spent longer in each fight.

Required 3.95 · optional 2.90

The required path is the hard path

Cross-referencing which of the 46 bosses Silksong actually forces you to fight, the difference is large and not subtle. The 15 mandatory bosses average 3.95. The 31 optional ones average 2.90 — a gap of more than a full point on a five-point scale (p = 0.0001).

Where the status comes fromMandatory / optional and act are taken from a published Silksong boss guide, not from the survey. Cited in the footer.

That inverts the usual expectation for a Metroidvania, where optional content is where the real challenge hides. In Silksong the hardest fights are mostly on the critical path, and the optional roster is where the walk-throughs live: Plasmified Zango, Palestag, Broodmother, both Savage Beastflies, Voltvyrm — all optional, all rated below 2.2.

Fig. 8 — Required and optional bosses, by how hard players found them

Required — 15 bosses
Optional — 31 bosses
Each mark is one boss. The required roster sits almost entirely in the upper half of the scale — its easiest member is Moss Mother, the tutorial boss. The optional roster spans the whole range, and contains both the easiest boss in the game and the hardest.

Because the exception at the top is a big one. First Sinner — the hardest boss in Silksong — is optional. It sits behind the bottom-most Heretic door in the Slab and has to be summoned deliberately. Only 1.3% of respondents hadn't fought it. Tormented Trobbio, eighth-hardest, is optional too.

Two of the top tenOf the ten hardest bosses, eight are mandatory. The two that aren't are the hardest fight in the game and the eighth-hardest.

So the shape of Silksong's difficulty is: a demanding critical path, a long tail of easy optional content, and a small number of hidden fights that are harder than anything the game requires. And this audience found nearly all of them.

What people actually missed

Skip rates confirm it from the other direction. Fourteen bosses were skipped by more than 5% of respondents, and they are overwhelmingly the easy optional ones — skip rate correlates negatively with difficulty (−0.25). Silksong gates its hidden content behind exploration, not skill.

Fig. 9 — What players miss is not what beats them

Act 1
Act 2
Act 3
Horizontal axis is the share who answered "Did not fight." The trend runs downward. The one boss far off to the right is a special case, explained below.

That outlier deserves its own note. Summoned Savior was never findable in a normal playthrough: it exists only in Steel Soul mode, Silksong's permadeath option, which unlocks after finishing the game once and entering the Konami code. Its 69.6% skip rate isn't players missing a hidden room — it's the share who never played permadeath.

30.4% did a permadeath runWhich is the single most telling statistic about this sample. The 4,967 who rated Summoned Savior had fought 45.2 bosses on average, against 42.2 for everyone else.

Read the other way, this is a striking measure of the audience: nearly a third of respondents had completed Silksong at least once and then started again on permadeath, far enough in to reach a secret boss in Moss Grotto. Its score of 2.28 comes from the most hardened subgroup in the entire sample, which is exactly why it should not be compared directly with the rest of the table.

Trobbio and Tormented Trobbio · 0.61

The data reconstructs the game's own structure

A useful test of whether a survey measured anything real: does it recover structure nobody told it about? Strip out each player's personal harshness, then correlate every pair of bosses. The strongest relationships in the entire 46×46 matrix are, without exception, the same boss fought twice.

Fig. 10 — Strongest rating correlations, after removing per-player severity

Every one of the top four pairs is a repeat encounter or a returning family. Nothing in the survey grouped these bosses — the columns were alphabetical within each act. The signal is strong enough to top all 1,035 possible pairs.

Below the repeat encounters, a second cluster appears with no shared name at all: First Sinner and Skarrsinger Karmelita (0.43), Lost Lace and Karmelita (0.38), First Sinner and Widow (0.38). That is the endgame gauntlet showing up as a statistical object — the fights that filter the same players out.

Average completion 93.7% · 19.0% fought all 46

Who actually answered this

Every number above carries the same caveat, and it is not a small one. The average respondent had fought 43.1 of 46 bosses. Nineteen percent had fought every one. Fewer than 5% had fought 35 or fewer.

Not the playerbaseThis is Mossbag's audience: players invested enough in Hollow Knight to watch lore analysis, and far enough into Silksong within weeks of release to rate its endgame.

Then there is the response curve. 78.3% of all 16,361 responses arrived in the first three days, 7,625 on 10 December alone. This is one viral spike, which means the sample is a snapshot of one community at one moment rather than a rolling cross-section of players.

Fig. 11 — The sample, two ways

Left: responses per day, 9 Dec – 5 Feb. Right: how many of the 46 bosses each respondent had fought. Both distributions are extremely lopsided — one toward a single day, the other toward near-total completion.

What this means for the numbers

Every difficulty score here is a lower bound for a general player. An audience this experienced compresses the low end of the scale: fights a first-timer would struggle with get scored 2 by someone on their fourth Hollow Knight playthrough. The ordering is likely robust to this. The absolute levels are not.

One more wrinkle sits in the timestamps. The 1,324 people who answered more than a week after launch did not score things the way the viral-day crowd did — they pushed the hard bosses higher and the easy bosses lower, spreading their answers wider rather than compressing them. The relationship between a boss's difficulty and how much late respondents shifted it is +0.44. Late arrivals polarise. Whether that is a different kind of player or a different kind of memory, the survey cannot say.

Order vs difficulty: r = 0.15, not significant

There is no difficulty curve

With the bosses placed in the order a player actually meets them, this becomes a testable claim rather than an impression. Correlating encounter position against difficulty across all 46 gives r = 0.15, p = 0.32. Restricted to the fifteen mandatory fights — the actual critical path — it is r = 0.37, p = 0.18. Neither is statistically distinguishable from no relationship at all.

Encounter orderThe canonical 1-to-48 boss order from a published guide, mapped onto the survey's 46 columns. Source in the footer.

Silksong does not ramp. It oscillates, and violently: Widow at position 10 scores 4.53, then Moss Mother II at 11 scores 2.39 and Savage Beastfly II at 12 scores 2.06. First Sinner, the hardest fight in the game, sits at position 26 of 48 — the middle. Broodmother, the fourth easiest, comes immediately after it.

Fig. 12 — Difficulty against the order players meet each boss

Required
Optional
Act bands shaded
Bosses in canonical encounter order, 1 to 48. The line links the mandatory path. There is no ascending trend — the sawtooth is the actual shape of the game. Hover any point for detail.

What does change across the game is the spread. Act 1 covers 2.48 points from its easiest boss to its hardest; Act 2 covers 2.92; Act 3 covers 3.13, holding both Skarrsinger Karmelita at 4.76 and Plasmified Zango at 1.63. Average difficulty drifts only slightly upward by act — 3.14, 3.32, 3.34 — while the share of maximum scores climbs from 39.3% to 46.9% to 47.2%.

Silksong doesn't get harder so much as it gets wider. The late game raises its ceiling without raising its floor, and the result is not mounting pressure but sharper and sharper contrast between the walls and the walk-throughs. A player's memory of "the difficulty spike" is really a memory of one of these teeth.

All 46 bosses · click any column to sort

The full table

Hard % and Easy % are the shares scoring 4–5 and 1–2. Divided is the 0–100 split score. Rank range is the 95% confidence range on the position from 1,000 resamples. Vet gap is the completionist difference, blank where the boss was skipped too often to compare fairly.

#BossRegion ActReqOrder Hard %Easy %Avg DividedRank range Missed %Vet gapAnswers