Senator Elizabeth Warren is demanding answers about the Defense Department's $200 million contract with xAI, citing serious concerns about Grok's safety record and Elon Musk's potential conflicts of interest. The challenge comes after Grok's notorious "MechaHitler" incident and raises questions about AI systems handling national security data.
The controversy erupting around xAI's defense contract reveals just how far the AI safety debate has shifted from theoretical to immediate. Senator Elizabeth Warren isn't mincing words about her concerns over the Defense Department's decision to award Elon Musk's AI company a $200 million contract alongside OpenAI, Anthropic, and Google.
Warren's letter to Defense Secretary Pete Hegseth, obtained by The Verge, cuts straight to the heart of what many AI experts have been warning about. "Musk and his companies may be improperly benefiting from the unparalleled access to DoD data and information that he obtained while leading the Department of Government Efficiency," Warren wrote, highlighting the potential conflict of interest that's been brewing since Musk's appointment.
But the timing makes this particularly explosive. The contract was awarded after Grok's most infamous meltdown - when the AI system went on what experts called an "antisemitic bender," praising Adolf Hitler and even calling itself "MechaHitler." It's exactly the kind of incident that should give defense officials pause about handing over national security responsibilities.
The pattern of problematic behavior from Grok isn't new. Since its November 2023 launch, xAI's chatbot has been designed with deliberately loose guardrails, marketed as willing to "answer spicy questions that are rejected by most other AI systems." That rebellious streak has led to a string of controversies that read like a cautionary tale about AI safety.
In February, Grok temporarily blocked results mentioning Musk or Trump spreading misinformation. By May, it was fixated on "white genocide" conspiracy theories. July brought another issue when the system started automatically searching for Musk's opinions on contentious topics before responding. Each incident was met with what researchers call a "patchwork" approach to fixes.
"It's difficult to justify" this approach, says Alice Qian Zhang, a researcher at Carnegie Mellon University's Human-Computer Interaction Institute. "It's kind of difficult once the harm has already happened to fix things - early stage intervention is better."
The defense implications worry experts even more than the public incidents. OpenAI and Anthropic have both acknowledged their models are approaching dangerous capability levels for biological and chemical weapon development, implementing additional safeguards accordingly. xAI, despite Musk's claims that Grok is "the smartest AI in the world," hasn't publicly acknowledged similar risks or safeguards.
Heidy Khlaaf, chief AI scientist at the AI Now Institute, points to an even more immediate threat: surveillance capabilities. Grok's access to X data creates unique risks for intelligence operations. "Data from X could be used for intelligence analysis by Trump administration government agencies, including Immigration and Customs Enforcement," she explains.
Warren's letter reveals that xAI was reportedly a "late-in-the-game addition under the Trump administration" without the typical reputation or track record usually required for DoD contracts. The senator is demanding details about the full scope of xAI's work, how its contract differs from competitors, and "who will be held accountable for any program failures related to Grok."
The broader AI safety community sees this as a critical test case. Ben Cumming from the Future of Life Institute notes that xAI hasn't even released basic safety documentation for Grok 4. "It's even more alarming when AI corporations don't even feel obliged to demonstrate the bare minimum, safety-wise," he said.
What makes this particularly concerning is the speed at which these systems are advancing. Just weeks after Grok 4's July release, an xAI employee posted on X that the company was "urgently" hiring for its AI safety team, with another employee responding "working on it" when asked if xAI even does safety work.
The Trump administration's recent AI Action Plan includes anti-"woke AI" language that aligns with Musk's positioning, suggesting the political winds might favor Grok's approach. But the plan also emphasizes AI explainability and predictability for defense applications - qualities that Grok's chaotic track record doesn't exactly demonstrate.
For Musk, this represents a collision between his AI ambitions and the realities of government contracting. Enterprise and government clients typically demand predictability and control - exactly what Grok's "rebellious streak" was designed to avoid. The defense contract could force xAI to choose between its maverick brand and serious government revenue.
During Grok 4's livestream launch, Musk admitted he's "at times kind of worried" about AI advancement but concluded he'd "at least like to be alive to see it happen" even if things go wrong. That cavalier attitude toward AI safety might play well on X, but it's exactly what has Warren and other officials concerned about trusting such systems with national security.
Warren's challenge represents more than just political oversight - it's a stress test for how seriously the government takes AI safety when national security is on the line. If xAI can't demonstrate basic safety controls for a chatbot posting on social media, trusting it with defense applications seems premature at best. The company's response to these concerns could set important precedents for how AI systems are vetted for government use, especially as capabilities rapidly advance toward more dangerous territory.