The Decision Failure Behind the AI Failure

Twenty prominent AI incidents reveal a common executive failure: uncertain machine output was allowed to acquire authority before evidence, limits and consequences had been properly tested.

LEADERSHIP & DECISION-MAKING

Dr Danie Adendorff

7/21/202615 min read

The Decision Failure Behind the AI Failure

What twenty documented incidents reveal about executive judgement before consequence

By Dr Danie Adendorff DSc (c.h), MSc

Artificial intelligence is usually blamed at the moment it produces the wrong answer. That is understandable. The visible error may be a fabricated legal authority, a discriminatory score, an unsafe vehicle movement, a false customer-service number or a convincing synthetic executive. Yet the machine output is only the beginning of the failure. Consequence arrives later, when an organisation decides that the output is credible enough, complete enough or authoritative enough to act upon.

A review of twenty documented incidents from 2016 to 2026 shows that the technical mechanisms varied widely. Microsoft Tay was manipulated in public. Amazon’s recruitment tool absorbed historical gender patterns. A healthcare algorithm used cost as a proxy for need. Zillow placed uncertain forecasts inside a capital-intensive acquisition process. Samsung employees entered confidential information into a public model. Lawyers and consultants allowed fabricated material to pass into courts and professional reports. Deepfakes defeated a payment process built around visual trust. Autonomous and driver-assistance systems converted uncertainty into physical movement.

These are not twenty versions of the same model defect. Nor are they all examples of equally poor management. McDonald’s, for example, used a pilot to discover operational limits and then ended that implementation. Honda’s braking investigation did not establish one common technical cause for every complaint. Arup was the target of an external criminal attack, not the operator of a defective internal AI system. The cases must therefore be handled with care.

“The decisive failure occurs when uncertain output is allowed to cross into publication, payment, employment, clinical advice, capital commitment or physical action without proportionate judgement.”

Even with those distinctions, the pattern is difficult to ignore. Most of the serious consequences were not produced by AI alone. They were created by an organisational decision architecture that defined the wrong problem, trusted a weak proxy, ignored an operating boundary, relied on ceremonial oversight, underestimated adversaries, failed to map reversibility or fragmented accountability. In Decision Before Consequence (DBC) terms, the model produced the signal; the organisation converted it into authority.

Twenty Different Incidents, a Small Number of Repeating Judgement Failures

The case literature can be reduced to seven recurring failures of judgement. They overlap, and several incidents contain more than one. That interaction matters: serious failures are rarely clean, single-cause events.

The business framed the wrong decision

Amazon’s experimental recruitment system did not discover an objective measure of merit. It learned from a historical environment in which male applicants and employees predominated, then treated similarity to that history as a signal of future suitability (Dastin, 2018). The Optum health-management algorithm made a related error at a deeper level. It predicted future healthcare expenditure when the real decision concerned medical need. Because Black patients with comparable illness had historically received less expenditure, the model systematically understated their need for additional support (Obermeyer et al., 2019).

In both cases, the model was not simply “biased” in the abstract. The organisation had defined the decision badly. Historical selection was confused with merit; expenditure was confused with illness. The data answered the question supplied to the system. The failure lay in choosing a question that was administratively convenient but substantively wrong.

iTutorGroup represents the blunt version of the same problem. The EEOC alleged that automated screening rejected women aged 55 or older and men aged 60 or older. Here, the harmful rule was not inferred through a subtle statistical proxy; it was allegedly built directly into the process (EEOC, 2023). Automation did not create the value judgement. It made that judgement faster, more consistent and less visible.

The organisation used a capability outside its proven boundary

Watson for Oncology carried the public image of a cognitive medical authority, but its recommendations reflected selected expert preferences, hypothetical cases and a limited validation base. Independent studies found uneven concordance with local clinical decisions (Ross and Swetlitz, 2018; Lee et al., 2018). Photo-based vehicle-damage systems could classify visible surface damage, but they could not inspect hidden structural or suspension problems. The failure came when a preliminary image-based estimate acquired the practical weight of a complete engineering assessment (Marshall, 2021).

Tessa, the National Eating Disorders Association chatbot, crossed a different boundary. Advice that might appear ordinary in a generic wellness application became hazardous when directed at people seeking eating-disorder support. Calorie restriction and weight-loss guidance were not merely inaccurate in context; they could reinforce the condition the system was supposed to address (Hoover, 2023).

McDonald’s automated drive-through pilot encountered the ordinary chaos of the physical world: traffic noise, passengers speaking at once, accents, corrections and incomplete orders. Ending the IBM implementation did not prove that voice ordering was impossible. It showed that benchmark capability had not translated into dependable end-to-end service under those conditions (Associated Press, 2024). In that respect, McDonald’s response was less a governance failure than a useful act of adaptation.

Prediction was allowed to become commitment

Zillow Offers is often compressed into the slogan that “AI lost US$881 million”. That account is too crude. The programme combined forecasting error, market volatility, labour and renovation constraints, inventory exposure and strategic scale. The judgement failure was more precise: uncertain price forecasts were connected to repeated purchases of illiquid assets. A modest average error became dangerous because the errors were correlated and the decisions were difficult to reverse (Zillow Group, 2021; 2022).

The same structural problem appears in physical systems. Uber’s developmental vehicle depended on a safety operator to intervene during rare automation failures, even though passive monitoring of a usually reliable system predictably encourages disengagement. The NTSB identified inadequate risk assessment, ineffective supervision and weak controls against automation complacency alongside the operator’s distraction (NTSB, 2019). Cruise’s post-collision system continued moving after the environment had become radically uncertain, and the company later faced a penalty for incomplete reporting to the regulator (NHTSA, 2024a).

Honda’s collision-mitigation investigation must be treated more cautiously because one common technical cause had not been established. The business-judgement issue is therefore not that Honda knowingly deployed one proven defective model. It is whether false-positive risk, field complaints, dealer escalation and investigation thresholds were governed rigorously enough as the pattern developed (NHTSA, 2024b).

When Output Becomes Authority

The most consistent failure across the later incidents was not poor prediction. It was misplaced authority. A system could generate language, retrieve information or imitate a familiar person. The organisation then permitted that output to function as if it were verified policy, legal authority, executive instruction or institutional knowledge.

Fluency was mistaken for evidence

In Couvrette v Wisnovsky, lawyers submitted briefs containing fifteen non-existent authorities and eight fabricated quotations. The language model created the initiating error, but qualified professionals converted it into a court filing. They did not retrieve the cases, check the quotations or stop when the material failed elementary verification. The human reviewer existed, but the review did not (Couvrette v Wisnovsky, 2025).

KPMG’s withdrawn agentic-AI report shows the same failure under institutional branding. Several named organisations disputed claims made about their AI activities. The precise role of generative AI in the production workflow was not publicly established, so it would be irresponsible to invent it. What is clear is that unsupported case material passed through a professional publication process and appeared under the authority of a major advisory firm (Financial Times, 2026).

Google AI Overviews extended the problem to mass public information. Traditional search ranks sources; a generative overview selects, combines and states a conclusion. A Munich court treated false summaries about publishers as Google’s own substantive statements rather than neutral links. The design had shifted Google from directing users towards evidence to publishing an answer, but the assurance model had not fully caught up with that change in role (Library of Congress, 2026).

The interface created authority the system did not possess

Air Canada’s chatbot gave incorrect advice about retrospective bereavement fares. The tribunal rejected the proposition that the chatbot was separate from the airline. It appeared on an official channel, represented company policy and influenced a customer’s decision. The liability remained with the organisation (Moffatt v Air Canada, 2024).

The Chevrolet dealership chatbot produced a more theatrical version of the same weakness. A user manipulated it into appearing to agree to a one-dollar sale. No car was transferred and no enforceable sale was established, but the incident exposed the underlying control failure: a conversational system placed on an official sales website could produce language resembling a binding commitment even though it had no legitimate pricing or contractual authority (Business Insider, 2023).

The poisoned customer-service number showed that retrieval can carry the same false authority. A fraudulent number circulated online and was presented by a Google AI Overview as trusted contact information. The model may have retrieved the number accurately from contaminated sources. That does not make the answer safe. Repetition across the web is not authentication, particularly when the query leads directly towards payment or account access (Ovide, 2025).

Synthetic presence defeated an old authentication model

The Arup fraud was not an internal AI malfunction. Criminals used synthetic video and audio to imitate executives and colleagues during a live conference. A finance employee authorised transfers totalling roughly US$25 million. The attack exploited a business process that treated familiar appearance, voice, hierarchy and group presence as sufficient confirmation of authority (Financial Times, 2024).

Blaming the employee for failing to “spot the deepfake” misses the strategic point. As synthetic media improve, visual recognition becomes a weaker control. A video call can carry an instruction; it should no longer authenticate a high-value transaction. The technology changed, but the organisation’s verification architecture remained anchored in an earlier information environment.

The Human Was Present, but Not Necessarily in Control

Several organisations could claim that a human remained in the loop. The phrase offered reassurance, but often concealed a weak control design. Human presence is not the same as human judgement. A person may lack time, attention, independent evidence, technical competence or the authority to stop the process.

Uber’s safety operator was physically present but was assigned a vigilance task that automation itself made difficult. The lawyers in Couvrette were professionally accountable but accepted fluent output without evidential review. KPMG’s report passed through human authorship and publication processes, yet disputed claims were not stopped. The Arup employee made a human decision, but inside a verification system that allowed synthetic presence to manufacture confidence. Air Canada’s customer could theoretically consult another webpage, but the official chatbot had already spoken with the airline’s authority.

This is automation bias in practical form. People tend to rely on automated systems when those systems usually perform well, appear confident or reduce cognitive effort. Parasuraman and Riley described misuse as inappropriate reliance on automation without adequate regard for its limits. Later work on trust in automation stressed calibration: reliance should match actual capability, not interface polish or institutional prestige (Parasuraman and Riley, 1997; Lee and See, 2004).

The twenty incidents suggest three recurring defects in human oversight. First, the intervention point came too late, after a vehicle had already committed to movement or after a transfer had already been authorised. Second, the human lacked independent evidence: the lawyer checked AI text against more AI text; the employee saw several deepfake colleagues supporting the same instruction; the customer saw an answer presented by the company’s own interface. Third, the human role had become ceremonial. Approval remained formally human, while the substance of judgement had already been outsourced.

A human return point is real only when the person has time, evidence, competence and power to stop the decision.

Scale, Adversaries and the Cost of Irreversibility

The same model error can be trivial or catastrophic depending on what follows it. A poor draft can be deleted. A house purchase, court filing, bank transfer, vehicle movement or public accusation may be far harder to reverse. The cases therefore have to be analysed not only by accuracy, but by coupling, scale and reversibility.

Zillow multiplied forecast uncertainty across a portfolio of illiquid assets. Uber and Cruise coupled perception and control to physical movement. Arup’s fraud moved money beyond the organisation’s immediate reach. Google and KPMG attached institutional authority to claims that could continue circulating after correction. Tessa addressed a vulnerable population in which harmful advice could reinforce dangerous behaviour before a clinician ever saw the exchange.

Perrow’s Normal Accidents theory is useful here. In complex, tightly coupled systems, unexpected interactions develop quickly and leave little room for correction. Not every AI application is a “normal accident” system, but the principle is relevant whenever components, users and decisions are closely connected. A small upstream error becomes a cascade when the organisation removes friction from the route to action (Perrow, 1999).

Adversarial pressure intensifies the problem. Tay was released into an environment containing coordinated hostile users. The Chevrolet chatbot was prompt-injected. Samsung’s public-model use exposed a route for unintended information egress. The Arup attackers manufactured identity. The fraudulent service number exploited a web environment that criminals could poison. These were not extraordinary edge cases. Public AI operates in a contested information space. Any business threat model that assumes cooperative users is already incomplete.

The adaptation record is mixed. Microsoft removed Tay quickly. Amazon abandoned the recruitment tool. McDonald’s ended its pilot rather than force a weak deployment into scale. NEDA suspended Tessa. Those decisions do not erase the original failures, but they demonstrate the value of stop rules. Cruise’s incomplete reporting illustrates the opposite: an organisation cannot learn responsibly when the evidence required to understand the incident is not fully disclosed.

What Decision Before Consequence Would Change

DBC does not replace technical AI safety. It cannot substitute for model testing, cybersecurity, fairness analysis, clinical validation, red teaming, NIST’s AI Risk Management Framework or ISO/IEC 42001. Nor does it make an inaccurate model accurate. Its contribution lies elsewhere: it governs the point at which machine output is converted into organisational judgement and action.

That distinction matters because the twenty incidents repeatedly crossed the same boundary. The model produced a score, forecast, sentence, image, recommendation or alert. The organisation decided what that output was allowed to mean.

The Executive Intelligence Pipeline

The Executive Intelligence Pipeline can be read directly against the case evidence.

At the signal stage, the organisation asks where the information came from and what might have contaminated it. Biased historical resumes, unequal healthcare expenditure, poisoned contact details, hostile public prompts and confidential source code all entered the pipeline without sufficient classification.

Validation asks whether the claim, source or system is reliable enough for the intended decision. Fabricated legal citations, unsupported KPMG case studies, inconsistent Air Canada policies and weakly generalised clinical recommendations should have failed here.

Interpretation asks what the output actually means. A forecast is not certainty. Expenditure is not illness. A photograph is not a structural inspection. A video call is not authentication. The DBC discipline is to make those assumptions explicit before the output is allowed to carry weight.

Escalation determines when the automated route must stop. Tessa should have escalated high-risk conversations. Hidden vehicle damage should have triggered physical inspection. Sensitive customer-policy questions should have moved to an authorised employee. Repeated anomalous braking complaints should have entered a structured safety review.

Decision and action define who may authorise the consequence. A chatbot may converse without contracting. A model may rank without rejecting. A valuation may inform without purchasing. A language model may draft without filing. A video call may communicate without authenticating a transfer.

Adaptation closes the loop. Near misses, disputed outputs and weak pilots are not embarrassments to conceal; they are intelligence about the operating system. A manipulated chatbot that does not complete a sale is still a warning. A clinician who rejects a poor recommendation has exposed a failure mode. A pilot that cannot handle field conditions has performed a useful function if the organisation has the discipline to stop.

The Human Return Point

DBC’s Human Return Point is not a ceremonial signature and not a person placed somewhere near the automation. It is a designed transfer of authority back to an accountable human before the consequence becomes irreversible.

For Arup, that point would have required out-of-band confirmation and independently verified beneficiary details before the transfer. For Couvrette, it would have required the signing lawyer to retrieve every authority from an authenticated legal database. For Zillow, it would have imposed capital limits and stop thresholds independent of model confidence. For Tessa, it would have forced abstention and clinical escalation. For autonomous vehicles, it means either timely human control or a fail-closed system state when uncertainty exceeds safe limits.

The Human Return Point must pass four tests: the person has enough time; has evidence independent of the model; understands the decision; and possesses real authority to stop it. If one of those conditions is absent, the organisation has retained human accountability in name but not in fact.

Could DBC Have Prevented the Twenty Failures?

The honest answer is not a universal yes. DBC is a decision-governance framework, not a guarantee against technical defect, criminal ingenuity or human error. Its likely effect differs across the incidents.

Failures that DBC could probably have prevented

The strongest prevention claim applies where a simple authority or verification gate was missing. Samsung’s confidential information should not have crossed into a public model. iTutorGroup’s discriminatory rule should not have entered production. The Chevrolet chatbot should never have possessed apparent contractual authority. Arup’s transfer should have required independent confirmation. Air Canada’s policy answer should have come from a controlled source. The fraudulent telephone number should not have been presented without authentication. The Couvrette citations should have been retrieved and checked. Google should have applied a higher threshold to reputational allegations. KPMG should have verified every named organisational case before publication. Tessa should have abstained and escalated rather than provide contraindicated advice.

In these examples, DBC would not need to solve an unsolved technical problem. It would need the organisation to stop, verify, constrain or refuse. The preventive leverage is therefore high.

Failures that DBC could have materially mitigated

Tay, Uber, Watson for Oncology, Amazon recruitment, the Optum algorithm, vehicle-damage estimation, Zillow and Cruise involved broader socio-technical limitations. DBC could not guarantee correct classification, perfect forecasting or flawless perception. It could have restricted scope, reduced decisional authority, introduced exposure limits, strengthened external validation and created earlier stop points.

The likely result would not always have been “no incident”. In some cases it would have been a smaller incident, a shorter exposure period, a reversible pilot or a failure caught before scale. That is still meaningful governance. Prevention is not the only standard; containment matters.

Failures where DBC mainly supports monitoring and adaptation

Honda and McDonald’s require restraint in the claim. Honda’s investigation did not establish one technical cause for every report, so DBC’s contribution would lie in complaint aggregation, threshold setting, transparent escalation and field monitoring. McDonald’s, by contrast, tested a system, discovered that it did not meet the operational requirement and ended the implementation. That is close to the behaviour DBC should encourage: bounded experimentation, explicit performance criteria and the willingness to stop before sunk cost becomes strategy.

The wider verdict is nevertheless firm. The twenty incidents were not twenty unrelated accidents. They were varied expressions of one recurring executive problem: organisations moved from AI capability to consequence without governing the conversion. Sometimes they chose the wrong target. Sometimes they trusted the wrong evidence. Sometimes they gave a system authority it had never earned. Sometimes the human returned too late, without independent information or without power.

The AI did not decide alone. The organisation decided what the AI was allowed to become.

That is why DBC is relevant. Its value is not the promise that AI will stop making mistakes. Its value is the discipline of asking, before action: What decision is being made? What evidence supports it? Where does the model’s competence end? Who may authorise the consequence? What is the cost of error? Can the action be reversed? What signal stops the process? Who remains accountable when the machine appears certain but the evidence is not?

Most of the reviewed consequences would probably have been prevented or materially reduced had those questions been converted into enforceable workflow gates rather than left as broad ethical intentions. The central DBC proposition is therefore supported by the case literature: the decisive risk is not that AI can be wrong. It is that an organisation will allow wrongness to acquire authority, travel at operational speed and become difficult to reverse before accountable judgement returns.

References

Associated Press (2024) ‘McDonald’s is ending its test run of AI-powered drive-thrus with IBM’, 18 June. Available at: https://apnews.com/article/mcdonalds-ai-drive-thru-ibm-bebc898363f2d550e1a0cd3c682fa234 (Accessed: 20 July 2026).

Business Insider (2023) ‘A car dealership added an AI chatbot to its site. Then all hell broke loose’, 18 December. Available at: https://www.businessinsider.com/car-dealership-chevrolet-chatbot-chatgpt-pranks-chevy-2023-12 (Accessed: 20 July 2026).

Couvrette v Wisnovsky et al. (2025) No. 1:21-cv-00157-CL, Opinion and Order, United States District Court for the District of Oregon, 12 December.

Dastin, J. (2018) ‘Amazon scraps secret AI recruiting tool that showed bias against women’, Reuters, 10 October.

EEOC (2023) ‘iTutorGroup to pay $365,000 to settle EEOC discriminatory hiring suit’, US Equal Employment Opportunity Commission, 11 September.

Financial Times (2024) ‘Arup lost $25mn in Hong Kong deepfake video conference scam’, 16 May.

Financial Times (2026) ‘KPMG report contained AI hallucinations on benefits of AI’, 12 June.

Hoover, A. (2023) ‘An eating disorder chatbot is suspended for giving harmful advice’, Wired, 1 June.

Lee, J.D. and See, K.A. (2004) ‘Trust in automation: Designing for appropriate reliance’, Human Factors, 46(1), pp. 50–80. doi: 10.1518/hfes.46.1.50_30392.

Lee, W.S. et al. (2018) ‘Assessing concordance with Watson for Oncology, a cognitive computing decision support system for colon cancer treatment in Korea’, JCO Clinical Cancer Informatics, 2, pp. 1–8. doi: 10.1200/CCI.17.00109.

Library of Congress (2026) ‘Germany: Court holds Google liable for incorrect AI Overviews’, Global Legal Monitor, 17 July.

Marshall, A. (2021) ‘AI comes to car repair, and body shop owners are not happy’, Wired, 13 April.

Microsoft (2016) ‘Learning from Tay’s introduction’, Official Microsoft Blog, 25 March.

Moffatt v Air Canada (2024) BCCRT 149, British Columbia Civil Resolution Tribunal, 14 February.

NHTSA (2024a) ‘GM’s Cruise failed to fully report pedestrian crash’, National Highway Traffic Safety Administration, 30 September.

NHTSA (2024b) Engineering Analysis EA24002: Inadvertent Activation of Collision Mitigation Braking System. Washington, DC: National Highway Traffic Safety Administration.

NTSB (2019) Collision Between Vehicle Controlled by Developmental Automated Driving System and Pedestrian, Tempe, Arizona, March 18, 2018. Highway Accident Report NTSB/HAR-19/03. Washington, DC: National Transportation Safety Board.

Obermeyer, Z., Powers, B., Vogeli, C. and Mullainathan, S. (2019) ‘Dissecting racial bias in an algorithm used to manage the health of populations’, Science, 366(6464), pp. 447–453. doi: 10.1126/science.aax2342.

Ovide, S. (2025) ‘Google’s AI pointed him to a customer service number. It was a scam’, The Washington Post, 15 August.

Parasuraman, R. and Riley, V. (1997) ‘Humans and automation: Use, misuse, disuse, abuse’, Human Factors, 39(2), pp. 230–253. doi: 10.1518/001872097778543886.

Perrow, C. (1999) Normal Accidents: Living with High-Risk Technologies. Princeton, NJ: Princeton University Press.

Ross, C. and Swetlitz, I. (2018) ‘IBM’s Watson recommended unsafe and incorrect cancer treatments, internal documents show’, STAT, 25 July.

Wilkinson, L. (2023) ‘Samsung employees leaked corporate data in ChatGPT: report’, CIO Dive, 7 April.

Zillow Group (2021) ‘Zillow Group reports third-quarter 2021 financial results and shares plan to wind down Zillow Offers operations’, 2 November.

Zillow Group (2022) ‘Zillow Group reports fourth-quarter and full-year 2021 financial results’, 10 February.

Author Workflow Disclosure

This article was researched and drafted with artificial-intelligence assistance under direct human instruction and editorial control. AI output was not treated as evidence. Material claims were checked against court decisions, regulatory records, company filings, peer-reviewed scholarship and established reporting. Analytical judgements, source selection and final publication responsibility remain with the author.