The most quoted number in DUI trial work is the eye test's 88 percent accuracy figure. It is real, and it comes from a study the State owns. The same table, on the same page, shows the test flagging 37 percent of the sober drivers who took it. Both numbers are correct, and the only difference is which question you ask. This episode walks the page both numbers come from, the laboratory the test was born in, and a validation study run with the Pinellas County Sheriff's Office.
It is made for defense lawyers preparing a motion to suppress or a cross examination, and it uses only documents the State already owns.
Episode 3
What this episode covers
- The San Diego long division: 290 drivers, 209 over the limit, 81 under, and the 30 wrong flags that make 88 percent and 37 percent the same arithmetic
- What the test's own authors say it was never designed to measure, and their word for the driving connection: irrelevant
- The 1977 laboratory finding that the eye test contributed almost nothing to the battery's own driving analogue
- The 1981 laboratory's 78 percent for an eyeballed angle against 88 for a measured one, and the officers the report said needed more practice
- The Pinellas County validation study: the 95 percent headline, its denominator, and the 51 of 57 refusals already scored a perfect six out of six
- Why the one test nobody can review always runs first, in NHTSA's own 1983 words
Watch this episode
Full transcript
There’s a number that decides a lot of these cases, and most people repeating it have never seen the page it comes from.
I hadn’t either, for a while. I’d heard it in court. I’d read it in a motion. I’d said it back to somebody without ever opening the report.
Then I opened the report.
I’m Rory Safir, and this is Reasonably Safir. Part two of five on horizontal gaze nystagmus.
Last time I showed you the test itself. The pen, the holds, the six clues. This time, the number that gets attached to it.
The number is eighty eight percent. And here’s the plan for the next ten minutes, so you always know where you are. First, the study that produced that number, and the second number sitting on the same page that nobody says out loud. Then where the test came from, because the laboratory that built it wrote down two things worth carrying into a hearing. Then the two studies that checked it later, one on people standing still, and one from my own county. And at the end, one question to ask, and why it works.
All right. The study.
San Diego, 1998. A field validation study run for the federal government by Jack Stuster and Marcelline Burns. Two hundred ninety drivers, stopped by real officers on real roads, each one with the eye test scored and the blood alcohol actually measured, so the two can be compared.
Here’s what the report says, and this part is real. Four or more clues means the subject is at or above a point zero eight, and using that criterion you’ll classify about eighty eight percent of your subjects accurately. That’s in the report. The officer who quotes it isn’t making it up.
Now stay on that same page with me. Same figure. Same four boxes.
Of those two hundred ninety drivers, two hundred nine were actually at or above the limit. Eighty one were actually under it.
Of those eighty one, thirty showed four or more clues.
Thirty out of eighty one. Thirty seven percent.
Let me say that again. The same table that gives the State its eighty eight percent shows the test flagging thirty seven percent of the sober drivers who took it.
Both numbers are correct. Same study, same data, and long division. The report divides the right calls by everybody and gets eighty eight. Divide the wrong flags by the sober drivers instead, and you get thirty seven. The only difference is which question you ask.
Nothing there is contested. Nothing came from a defense expert. That’s the State’s study, the State’s counts, and arithmetic.
Now let me be fair about it, because if you’re not fair about it you’ll get corrected in front of a jury. Those two numbers are not in conflict. A test can be right about most of the population it gets used on and still wrongly flag a third of the sober people who take it. They answer different questions. In that San Diego sample, roughly seventy two percent of the drivers were over the limit before anybody looked at anyone’s eyes. Run any test on a group that’s mostly guilty already and it will look accurate.
Which is the honest frame for the whole episode. The test performs well on a population already selected for drinking. Your client is asking a different question, which is how it performs on him.
And while we’re on that report, here’s the part that took me longest to really get, and the part I most want you to keep.
These tests were never validated to measure impairment. They were validated to predict a number.
Don’t take that from me. Take it from Stuster and Burns, in the same report. They write that many people, including some judges, believe the purpose of a field sobriety test is to measure driving impairment, and that they expect the tests to look like driving. They say horizontal gaze nystagmus lacks that kind of face validity, because it doesn’t appear to be linked to the requirements of driving a car. And then they say the reasoning is correct, but the assumption behind it is wrong, because the tests were never designed to measure driving impairment at all.
And then this. HGN’s apparent lack of face validity to driving tasks is irrelevant, because the objective of the test is to discriminate between drivers above and below the statutory limit, not to measure driving impairment.
Irrelevant. Their word, not mine.
So when the officer tells your jury the eyes showed impairment, he’s describing the test as something its own authors say it isn’t.
That’s the study everybody quotes. Now rewind twenty years, to the laboratory where this test was born. Two findings, and the State quotes neither of them.
The first is about driving. In the original 1977 work, the researchers connected the battery to a divided attention driving simulator task in the laboratory. All three tests together explained about a third of the variation in performance on that task.
And the eye test contributed almost nothing to that.
Almost nothing. On the closest thing to driving anyone ever tested it against.
The second is about precision. In the 1981 laboratory study, sorting people against a point one zero blood alcohol level, and I’m saying that threshold out loud because it is not the point zero eight from San Diego, the officers’ eyeballed angle of onset, used by itself, correctly classified seventy eight percent of subjects. When a machine measured the same angle instead, it was eighty eight.
That’s their own research telling you what ten points of accuracy get thrown away for the convenience of not carrying the device. He’s estimating forty five degrees on the side of a road with nothing but his arm.
Same laboratory, one more finding. They measured how well the officers agreed with each other, which is a different question from whether they were right. On counting clues for the physical tests, they agreed reasonably well. But the further the measure moved from counting toward judgment, the worse the agreement got, and by the second session, agreement on the arrest decision, the only output that matters to a human being, was the lowest thing they measured.
And the report says it in its own voice. The interrater reliability for the nystagmus score is not as high as expected, suggesting that the officers would profit from further training and practice with nystagmus.
That’s the study that built this test saying the officers needed more practice at it.
So that’s the number, and that’s the laboratory it came from. Now the two checks that came later.
The first comes from a study NHTSA itself cites in the current manual, so nobody can wave it off as defense literature. A 2003 paper in the journal Optometry, testing whether this works on a person standing, sitting, or lying down. Standing, at a point zero eight threshold, they measured a false alarm rate around thirty seven percent. Practically the same number as the San Diego long division, from a research group independent of the NHTSA validation team.
The same paper reports test retest reliability for standing HGN at point five eight nine, against its own stated benchmark that a highly reliable test of this kind runs around point seven.
And now I’m going to argue against myself, because you need this before you use it. Those authors do not conclude what I just implied. They conclude the opposite. They write that the test is highly reliable, and they say that point five eight nine is not statistically different from the other reliability figures in their own study. Their subjects were also volunteers at police training workshops, tested indoors, not drivers stopped on a highway.
So do not walk into a hearing and tell a judge this paper says the test fails. It doesn’t. What it gives you is a number the paper printed, next to a benchmark the paper also printed, and the honest sentence is that the standing figure sits below the benchmark the authors themselves named. Say that much and no more, because the State has read the conclusion too.
The second check is close to home. In 1997 Dr. Marcelline Burns ran a validation study right here in Pinellas County, with the Pinellas County Sheriff’s Office. My county.
The headline is that more than ninety five percent of the deputies’ arrest decisions were correct. And I want to say the denominator out loud, because this episode would be a fraud otherwise. That’s arrests only. A hundred ninety seven out of two hundred six. Their release decisions in the same study were eighty two percent correct, and the report notes the officers erred more often by releasing an impaired driver than by arresting a sober one.
Two things underneath it.
Fifty seven drivers in that study refused to give a breath sample. Fifty one of them had already been scored a perfect six clues out of six. The handful left were either right at the arrest cutoff or had no eye score recorded at all.
And of the nine arrests the study counted as errors, meaning the driver was actually under the limit, in six of those nine the officer had recorded the maximum number of clues. Let me be fair about that the same way I was fair about San Diego. Five of the nine blew between point zero six three and point zero seven, and the report points out those drivers were still prosecutable under Florida’s presumptive statute. At least two were suspected drug cases, which the report offers as the reason for the high eye scores.
But sit with the shape of it. The maximum score, on people who were under the limit.
That’s what a test looks like when it’s sitting at the ceiling. It isn’t discriminating anymore. It’s confirming.
Last section, and it’s the reason all of this matters more than it looks.
If you’ve watched enough of these stops, you’ve noticed the eyes always go first. Before the walk and turn, before the one leg stand. And nobody gets arrested right after the eyes. He does the eyes, then the physical ones, and then he decides.
Which sounds like it makes the eye test less important. It makes it more important. And the reason isn’t a conspiracy theory. It’s how people work.
He runs the eyes at minute one and forms an impression. Then he grades the walk and turn at minute six. And by then he isn’t evaluating the walk and turn on its own terms anymore. He’s checking it against a conclusion he already has.
You don’t have to take that from me either. NHTSA’s own field researchers named it in 1983. They were explaining why you shouldn’t compare the three tests’ accuracy figures against each other, and they wrote that in most cases all three were given in the same order, with gaze nystagmus first, that the results of the gaze nystagmus test were then known to the officer, and that this may have had some subtle influence on his expectations and his scoring of the next two tests.
Their study. Their words. 1983.
So it isn’t one of three tests. It’s the anchor. The unreviewable one runs first and sets the frame for the two you can actually watch.
And now put that next to something in the officers’ reference material. There’s a passage about what to do when the clues don’t come out even. It acknowledges a clue can show up in one eye and not the other. And then it explains that most circumstances in which this happens involve the officer missing the observation of the supposedly absent clue. And it reminds him he can recheck any individual clue as needed.
So he arrives at the eyes already suspecting alcohol. The material tells him an uneven result probably means he missed something. And it tells him he can look again.
That is not a measurement. That’s a search.
So here’s where we’ve been. One study, two true numbers, eighty eight and thirty seven, and which one gets said in your client’s case depends entirely on who is doing the talking. A laboratory that watched the eye test contribute almost nothing to its own driving task. A standing test sitting below its own paper’s benchmark. And a study in my county where the ceiling score kept landing on people who were under the limit.
Your question for the hearing. Ask him what the eighty eight percent was measured against. Not whether he knows the figure. What it was measured against.
I’ll tell you now that most of them can’t answer it. And the reason isn’t that they’re hiding it from you.
Which is the next part. Three days in a classroom, and what nobody in that room ever gets asked.
I’m Rory Safir. If you want the manual editions or the studies from this episode, email me at Rory at the Safir lawyer dot com, and I’ll send them to you.
Take care of yourself, and take care of your clients.
Keep going
Part three, Three Days in a Classroom, goes inside the three day certification classroom. Part one, The Eyes, covers the test itself, and the pages on horizontal gaze nystagmus and why the eye test is treated as scientific evidence in Florida cover the same ground in writing.
Nothing on this page or in any episode is legal advice, and listening to it does not make me your lawyer.




