The neck flexor endurance test measures how long a patient can hold a small chin tuck and head lift against gravity, and a short hold time flags deep neck flexor weakness tied to neck pain, headache, or postural dysfunction. Clinicians perform it in supine, timing to fatigue or loss of form. It works well as a fast outpatient screen and pairs naturally with the craniocervical flexion test when motor control, not just endurance, needs a closer look.
TL;DR:
- The test is influenced by consistent positioning, two-finger lift reference, and strict stopping criteria to ensure reliable results.
- An inter-rater reliability of 0.66 indicates moderate consistency, so changes of at least ten to fifteen seconds should be considered meaningful.
- Low scores often result from superficial muscle substitution, highlighting the need to assess both endurance and motor control with complementary tests.
- Use the test cautiously in cases of cervical instability, recent surgery, dizziness, or severe neurological signs, and always compare progress to the patient’s baseline.
Table of Contents
- Neck Flexor Endurance Test Protocol: Setup and Execution
- Neck Flexor Endurance Test Results: Norms and What They Mean
- How Reliable Is the Neck Flexor Endurance Test?
- Neck Flexor Endurance Test vs. Craniocervical Flexion Test
- When to Use, Modify, or Skip This Test
- Getting Consistent, Trustworthy Results Every Time
- Integrating the Test Into a Full Posture-Retraining Plan
- Sources
Neck Flexor Endurance Test Protocol: Setup and Execution
Getting reproducible numbers from this test comes down to two things: consistent positioning and a stopwatch that starts the moment the chin tuck breaks form, not a moment later. Skip either one and your “38 seconds” from last month means nothing next to today’s “41 seconds.”
Setup. You need a treatment table and a stopwatch, nothing more. That low equipment bar is part of why the APTA lists this test as a practical option for patients with neck pain, headache, and suspected postural dysfunction, even in clinics without pressure biofeedback units. Position the patient supine, hook lying, with knees bent and feet flat. Arms rest at the sides. The head starts in a neutral, comfortable position on the table, not tipped back into extension.
The chin tuck cue. Ask the patient to nod the chin toward the chest, like they’re making a double chin, without actually flexing the neck up. Most patients get this wrong on the first try by jutting the jaw or tensing the whole neck. A useful verbal cue: “Nod your head yes, small and slow, then hold that nod while you lift.”
Verifying lift height. Once the chin tuck holds, the patient lifts the head roughly one inch off the table, about two finger-widths. Many clinicians place two fingers flat under the occiput before the lift starts. That gives a physical reference point: if the head sinks back down and touches the fingers, the trial is over. This two-finger technique also keeps lift height consistent across trials and across different clinicians testing the same patient later.
Here’s the step-by-step sequence to run in the room:
- Position the patient supine in hook lying with a neutral head start position.
- Cue the chin tuck and confirm it visually, checking for a visible fold or crease at the front of the neck rather than a stiff, braced jaw.
- Instruct the patient to lift the head about one inch, keeping the chin tucked throughout.
- Start the stopwatch the instant the head clears the table.
- Watch continuously for the stopping criteria below and stop the clock the moment one appears.
- Record the time, rest the patient for three minutes, then repeat for a second trial.
- Record the best of the two trials as the test score.
Stopping criteria matter as much as the setup, because a test with vague endpoints produces numbers nobody can trust. End the trial when any of these happen: the chin tuck breaks down, visible as separation or loss of that skin fold at the throat; the head touches the clinician’s fingers (or the table) for more than one second; the patient starts substituting with obvious neck flexion instead of holding the tucked position; or the patient stops voluntarily due to pain or fatigue. These pragmatic stop rules show up across most published protocols precisely because they’re easy to apply consistently, even for a clinician running the test for the first time.
Two trials with a three-minute rest between them is the standard recommendation. Skipping the rest, or cutting it to sixty seconds because the schedule is tight, will shorten the second trial’s hold time and make you think the patient fatigued more than they actually did.
Pro Tip: Watch the sternocleidomastoid, not just the chin. If the SCM cords up visibly at the front of the neck while the chin tuck looks fine, the patient has likely shifted the work to superficial flexors instead of the deep neck flexors you’re trying to test. That substitution pattern is the single most common reason for an inflated hold time.
Chart the raw seconds for both trials, not just the best one. A big gap between trial one and trial two (say, 45 seconds then 22 seconds) tells you something different than two consistent scores near each other, and that gap is worth a note in the record.
Neck Flexor Endurance Test Results: Norms and What They Mean
A healthy adult’s hold time on this test varies more than most clinicians expect, and that variability is exactly why a single universal cutoff doesn’t work well in practice.
The most frequently cited normative dataset comes from Doménech et al., who tested asymptomatic adults and reported a mean hold time of 38.9 ± 20.1 seconds for men and 29.4 ± 13.7 seconds for women. That gap between sexes is consistent, not incidental. Men tend to hold longer, and testing a female patient against a male-derived benchmark will make her look weaker than she is relative to her own population.
Clinical samples paint a different picture. One study of patients with neck pain found a mean neck flexor muscle endurance score of 44.9 ± 25.3 seconds, which sits inside the same broad range as the asymptomatic norms above. That overlap is a genuinely useful, if slightly uncomfortable, finding: hold time alone doesn’t cleanly separate people with neck pain from people without it. A patient can report significant pain and disability while still posting a respectable endurance score, and vice versa.
| Population | Mean hold time | Standard deviation | Source |
|---|---|---|---|
| Asymptomatic men | 38.9 s | ± 20.1 s | Doménech et al. |
| Asymptomatic women | 29.4 s | ± 13.7 s | Doménech et al. |
| Neck pain patients (mixed sex) | 44.9 s | ± 25.3 s | Clinical sample study |
The standard deviations here are wide relative to the means, some approaching half the average value. That spread reflects genuine, well-documented intersubject variability, not measurement sloppiness. Age and general activity level often show limited influence on hold time in these samples, which runs against the intuitive assumption that a fit, active patient will automatically test well. Endurance in the deep neck flexors behaves as something closer to a specific, trainable trait than a byproduct of overall fitness.
Given that spread, here’s how to use these numbers without overreaching:
- Treat scores well below the lower end of the normative range (roughly under 20 seconds for women, under 25 for men) as a flag worth investigating further, not a diagnosis.
- Use the patient’s own baseline as the real benchmark for progress. A jump from 15 seconds to 28 seconds over six weeks means more clinically than where that patient sits against a population average.
- Cross-reference low scores with symptom presentation. A short hold time paired with headache, upper trapezius overactivity, or a forward head posture pattern is a stronger clinical signal than a low number in isolation.
- Set patient-specific goals in small increments, five to ten second targets, rather than aiming straight for a normative mean that may not fit that patient’s age, sex, or baseline conditioning.
- Recheck at intervals of four to six weeks rather than week to week, since normal day-to-day variation can otherwise look like progress or regression that isn’t real.
Pair the raw hold time with other measures before drawing conclusions. A cervical extensor endurance test, a pain scale, and a disability index like the Neck Disability Index round out the picture, especially since flexor and extensor endurance correlate only moderately with each other, discussed further below. Relying on the flexor endurance score by itself, especially with a patient near the borderline of normative ranges, risks over interpreting one noisy number.
How Reliable Is the Neck Flexor Endurance Test?
Reliability numbers for this test sit in the “usable but not perfect” zone, and knowing exactly where the uncertainty lives changes how much weight you put on any single session’s result.
Doménech and colleagues reported an inter-rater intraclass correlation coefficient of 0.66, with a confidence interval spanning 0.34 to 0.86. That’s moderate reliability at best. An ICC of 0.66 means two different clinicians testing the same patient on the same day will often land on meaningfully different numbers, not identical ones. The wide confidence interval matters just as much as the point estimate: a true value could sit as low as 0.34, which is poor agreement, or as high as 0.86, which is good agreement. The honest takeaway is that this test’s reproducibility across raters isn’t nailed down tightly, and standardizing your cueing and stopping criteria (the two-finger technique, the same verbal script every time) is what pulls your own clinical practice toward the better end of that range.
By the numbers: an ICC of 0.66 falls into the range most reliability researchers label “moderate” agreement, meaningfully short of the 0.90 threshold generally expected for a measure used to make individual treatment decisions on its own.
On validity, the picture is similarly mixed rather than clean. A moderate correlation (r = 0.52, p = 0.003) exists between cervical flexor and extensor endurance in patients with neck pain, meaning the two capacities relate but don’t move in lockstep. A patient with poor flexor endurance won’t automatically have poor extensor endurance, and testing only one side of that pair leaves half the picture unexamined.
A few practical implications follow from this evidence:
- Don’t treat a single test session as gospel. Because inter-rater reliability is moderate, a change of a few seconds between two different clinicians’ assessments may reflect testing variation rather than real physical change.
- Look for changes of at least ten to fifteen seconds before calling something a meaningful clinical improvement, given the natural noise in the measurement.
- Test flexor and extensor endurance as a pair when possible, since the moderate correlation between them means neither one substitutes for the other.
- Keep the same clinician testing the same patient across a course of care whenever feasible, since that removes inter-rater variability from the equation entirely.
The existing evidence base also leans on relatively modest sample sizes and specific patient populations, so treat normative comparisons as directional guidance rather than fixed thresholds that apply identically to every patient walking into your clinic.
Neck Flexor Endurance Test vs. Craniocervical Flexion Test
These two tests measure genuinely different things, and confusing them leads to picking the wrong tool for the clinical question in front of you.
The craniocervical flexion test uses a pressure biofeedback unit placed under the neck’s natural curve. The patient performs a small, precise nodding motion, and the device tracks incremental pressure increases to assess neuromotor control and activation quality of the deep cervical flexors, longus colli and longus capitis specifically. It’s testing coordination and isolated activation, not raw stamina.
The head-lift endurance test covered throughout this article measures something closer to gross muscular endurance: how long the flexor group as a whole, deep and superficial, can sustain a loaded hold against gravity. These two constructs diverge often enough that a patient can perform reasonably well on one and poorly on the other. A patient with good motor control but low general conditioning might struggle on the endurance test while scoring fine on the CCFT, and the reverse pattern shows up just as often, particularly in patients who’ve learned to compensate with superficial muscles.
Here’s a workflow that uses both tests without wasting clinic time:
- Start with the endurance test as a quick screen. It requires no equipment beyond a stopwatch and takes under ten minutes including rest between trials.
- If the hold time is low, or if you observe early substitution with the sternocleidomastoid or scalenes, follow up with the CCFT to check whether the issue is motor control, raw endurance, or both.
- If a pressure biofeedback unit isn’t available, watching closely for that early SCM substitution during the endurance test itself gives an informal, lower-resolution proxy for motor control quality.
- Use the CCFT result to guide exercise selection: patients with poor activation quality generally need low-load motor control drills before they progress to endurance-focused holds.
- Reassess with whichever test flagged the original deficit, since that’s the more sensitive measure for tracking that specific patient’s progress.
Running both tests on every patient isn’t necessary. Reserve the CCFT for cases where the endurance test result is ambiguous, where headache or cervicogenic symptoms suggest a motor control component, or where initial exercise progressions aren’t producing the expected gains.
When to Use, Modify, or Skip This Test
The neck flexor endurance test earns its place with patients presenting with mechanical neck pain, cervicogenic headache, or postural complaints tied to prolonged sitting and screen use, since these groups make up the bulk of the evidence behind the normative data. It’s also a reasonable screening tool when a patient reports vague neck fatigue that worsens through the day, a pattern often linked to deep flexor endurance deficits rather than acute injury.
Certain presentations call for caution or an outright pass on this test:
- Hold off if you suspect upper cervical instability, particularly in patients with rheumatoid arthritis, Down syndrome, or a history of trauma affecting the craniovertebral junction, since sustained flexion loading is inappropriate until instability is ruled out.
- Avoid testing within the standard postoperative window after cervical spine surgery until the surgical team clears active loading.
- Screen for severe dizziness or vertigo before testing, since the supine position and sustained head lift can provoke symptoms in patients with vestibular involvement.
- Stop or defer testing if the patient shows uncontrolled radicular signs, meaning progressive neurological symptoms down the arm, since these need medical workup before a strength or endurance assessment adds value.
- Adapt the test for patients who can’t tolerate supine positioning for other reasons, such as late-stage pregnancy or severe reflux, by considering a modified seated protocol or postponing until tolerance improves, and note the modification clearly in the chart since it breaks comparability with standard norms.
For elderly patients, expect naturally shorter hold times and don’t force a comparison against a young adult normative mean. For hypermobile patients, watch closely for excessive cervical flexion substituting for a true chin tuck, since hypermobility often masks poor deep flexor activation behind an exaggerated but low-quality range of motion.
Results from this test should shape what you prescribe next, not just what you write in the chart. A low score with clean form points toward progressive endurance work, isometric holds building in duration over weeks. A low score with early substitution points toward motor control retraining first, isolating the deep flexors before loading them with a timed hold.
Getting Consistent, Trustworthy Results Every Time
Small technique errors account for most of the inconsistency clinicians run into with this test, and nearly all of them are fixable with a five-minute habit change.
The two-finger technique deserves a permanent spot in your routine. Placing two fingers under the occiput before the lift gives both you and the patient a physical reference for the roughly one-inch target, and it makes the “head touches down for more than one second” stopping rule easy to apply without guesswork. Pair that with a consistent verbal script: cue the chin tuck first, confirm it visually, then cue the lift, in that order, every single time. Changing your wording session to session introduces variability you don’t need.
The most common errors worth watching for:
- Inconsistent lift height between trials, usually from eyeballing it instead of using the two-finger reference.
- Rushing the rest interval below the recommended three minutes, which artificially shortens the second trial.
- Missing early SCM substitution because attention stays fixed on the chin instead of scanning the whole neck.
- Failing to document both trial times, which erases useful information about fatigue and consistency.
Pro Tip: Chart four things every time: both trial times, any observed compensation pattern, the patient’s pain response during the hold, and whether you’re following up with the CCFT. That four-point note takes fifteen seconds to write and saves real time on the next visit when you’re trying to remember what actually happened.
For patients who score low, start conservative. Isometric holds in the tucked position, working up in five-second increments, build endurance capacity without overloading a weak system on day one. Once form holds steady across a full set, progressive endurance work and cervical retraction drills round out a reasonable early-phase program.
Integrating the Test Into a Full Posture-Retraining Plan
A hold time in seconds tells you something real, but it’s still just one data point. The neck flexor endurance test should be treated as a starting measurement, not a finish line, because the number only matters in the context of what a patient’s neck does for the other twenty three and a half hours of the day.
Forward head posture develops slowly, usually across years of screen time, driving, and desk work, and it rarely resolves because someone did thirty seconds of chin tucks in a clinic once a week. The test result is useful precisely because it gives clinicians and informed patients alike an objective marker to return to. Did six weeks of targeted work move a 22 second hold to 35 seconds? That’s a number you can point to, distinct from a patient’s subjective sense that their neck “feels a little better.”
Our approach leans on education first: understanding why the deep neck flexors go quiet under chronic forward head positioning, and why strengthening them without addressing the postural habit that weakened them in the first place tends to produce short-lived gains. A clear picture of neck alignment paired with an honest endurance score gives a far more complete starting point than either measure alone.
For clinicians using this protocol with patients already working through posture correction, the test result should inform, not replace, a broader look at how someone sits, sleeps, and holds their phone. A deeper look at deep neck flexor anatomy explains why these particular muscles fatigue first and fail quietly, often long before a patient notices pain. We built our library of guides around that idea: measurement matters, but only when it’s paired with the habit-level changes that keep the gains from evaporating in a month.
Clinicians and patients working through a rehab plan are welcome to use our resources as a companion to whatever protocol is already in place in the clinic.
— Madhukar
Sources
For deeper detail beyond what’s covered here, the original Doménech et al. normative study remains the primary source for age and sex-based hold time benchmarks and reliability data. The APTA’s test summary page offers a concise professional reference for indications and scoring. For the motor-control side of assessment, Physiopedia’s CCFT overview explains the pressure biofeedback protocol in detail, and the cervical flexor-extensor correlation study is worth reading for anyone testing symptomatic populations rather than healthy volunteers.
This article is general information, not a substitute for advice from a qualified doctor. Consult a qualified healthcare professional about your own circumstances before acting on anything here.
- The Deep Neck Flexor Endurance Test: Normative Data Scores in Healthy Adults
- The relationship between cervical flexor endurance, cervical extensor endurance, VAS, and disability in subjects with neck pain
- Neck Flexor Muscle Endurance Test | APTA
