In July 2018 the American Civil Liberties Union of Northern California spent $12.33 of cloud credit to demonstrate something that had until then lived mostly in academic papers: a commercial face-matching service sold to police departments, run at the settings Amazon shipped it with, will confidently identify innocent people as arrestees. Comparing official photographs of every sitting US senator and representative against a database of 25,000 publicly available arrest photos, Amazon Rekognition returned 28 false matches, and those errors fell disproportionately on lawmakers of color. No one was arrested and no one was hurt, but the test converted a research finding into a political fact and set the terms of the American face-recognition fight for years afterward.

What the test did
The ACLU built a Rekognition face collection from 25,000 arrest photos, then searched it using public photos of all 535 members of Congress, leaving Amazon's default match settings in place. The software flagged 28 sitting legislators as matches for people who had been booked on criminal charges. The false matches spanned both parties, both chambers, and every age bracket, and included six members of the Congressional Black Caucus among them Representative John Lewis of Georgia, as the ACLU set out in its own account of the experiment. By the ACLU's count, roughly 40 percent of the erroneous matches involved people of color, who then made up about a fifth of Congress.
The organisation framed the exercise deliberately as a replication of real police practice rather than a laboratory curiosity: matching faces against booking photos was exactly the workflow Amazon and an Oregon sheriff's office had jointly documented publicly a year earlier, and the ACLU noted that the same office had already begun running such searches without local debate.

Independent checking and immediate fallout
The result was not left to rest on the ACLU's word alone. BuzzFeed News reported that Joshua Kroll, then a postdoctoral scholar at the University of California, Berkeley's School of Information, independently verified the test, and that he considered the outcome consistent with the wider machine-learning literature showing degraded performance on women and darker-skinned faces. The same day, Representatives John Lewis and Jimmy Gomez wrote to Jeff Bezos seeking a meeting about the misidentifications and the technology's effect on communities of colour.
Fortune and other outlets picked the finding up within hours, folding it into an existing campaign in which Amazon employees, shareholders and dozens of civil rights groups had already pressed the company to stop selling Rekognition to government agencies.

Amazon's rebuttal, and the shifting threshold
Amazon Web Services first told reporters the test would have looked different under its recommended practices, then, the following day, published a fuller response from machine-learning executive Matt Wood. Wood accepted the raw arithmetic, 28 wrong matches out of 535 at an 80 percent confidence setting, but argued the setting itself was the error, and reported that AWS had re-run a similar search against a corpus of more than 850,000 faces at a 99 percent threshold and seen no misidentifications at all.
The 80% confidence threshold used by the ACLU is far too low to ensure the accurate identification of individuals
Matt Wood, Amazon Web Services
Wood also stressed that in law enforcement deployments Rekognition was, in Amazon's description, used to narrow a field of candidates for human review rather than to make decisions on its own, and he compared discarding the technology over a badly chosen setting to throwing out an oven because it can burn a pizza, as GeekWire reported. The ACLU's answer was that Amazon's advice had moved faster than its product: the company's shipped default was 80 percent, its own 2017 joint post with the Oregon sheriff's office had used 85 percent, and within two days its public recommendation had climbed to 95 and then 99 percent.
In its five stages of grief over its dangerous face surveillance product, Amazon is clearly stuck at denial.
Jacob Snow, ACLU Foundation of Northern California

Timeline
- Jun 2017Amazon Web Services and Washington County, Oregon publish a joint post showing how to use Rekognition to search arrest photos, using an 85 percent confidence threshold, according to the ACLU's later timeline of events.
- May 2018The ACLU publishes an investigation reporting that Amazon is marketing Rekognition to law enforcement agencies, including Washington County and Orlando.
- 18 Jun 2018A petition with more than 150,000 signatures and a coalition letter from nearly 70 organisations are delivered to Amazon's Seattle headquarters; 19 investor groups send a shareholder letter to Jeff Bezos.
- 21 Jun 2018Amazon employees write to Jeff Bezos urging the company to stop selling Rekognition to police.
- 26 Jul 2018The ACLU of Northern California publishes its Rekognition test: 28 members of Congress falsely matched against a 25,000-image arrest photo database at Amazon's default settings, for $12.33 of compute. Reps. John Lewis and Jimmy Gomez write to Jeff Bezos the same day.
- 26 Jul 2018An AWS spokesperson tells reporters the results reflect an inappropriate confidence threshold and points to a 95 percent figure for law enforcement use.
- 27 Jul 2018AWS machine-learning executive Matt Wood publishes a blog post confirming the 28-of-535 figure at 80 percent confidence, reporting a zero misidentification rate in Amazon's own 99 percent threshold re-run, and recommending 99 percent for law enforcement.
- 27 Jul 2018ACLU attorney Jacob Snow responds that Amazon's recommended threshold moved from its own 80 percent default to 95 and then 99 percent within 48 hours, and renews the call for a congressional moratorium.
Why it moves the needle
The substantive dispute here is narrower than the headlines suggested, and more damning for it. Amazon did not claim the ACLU had faked anything; it claimed the customer had used the wrong knob. That is the point. A face-surveillance service was being marketed to police forces with a default configuration that its own vendor, once challenged, said was unfit for public-safety work, and at least one sheriff's office had been shown how to run mugshot searches at a threshold well below the 99 percent Amazon retroactively insisted upon. The failure mode was not a rogue algorithm but a deployed product behaving exactly as configured, with the burden of getting the configuration right pushed onto agencies whose errors land on arrestees rather than on shareholders.
Two features made this episode unusually durable. First, the victims of the false matches were legislators, which turned a measurable accuracy problem into a personal one for the people who write surveillance law. Second, the demonstration was cheap and trivially reproducible, so it could not be dismissed as a one-off artefact of a hostile researcher's setup. Bias in commercial face recognition had been documented earlier in 2018 in peer-reviewed work the ACLU itself cited; what this test added was a concrete, dateable, politically legible instance of a live product sold for policing producing racially skewed false accusations at its factory settings.