7 AI Agents Tasked with Classifying and Correcting Bitcoin Flaws: What Did They Find?
- The models Qwen 3.8, GPT 5.6, and Fable failed: one got stuck and two refused to work.
- 4 out of the 7 tested models are open-source and can be used without filters.
Paul Miller, a developer who maintains cryptography libraries used in Bitcoin and Ethereum projects, tested 7 artificial intelligence (AI) agents to classify and correct 5 real security vulnerabilities reported by the Bitcoin Red Team. The results showed marked differences between the models, both in the dollar cost incurred to run each one and in the accuracy of their diagnoses.
Of the 7 agents evaluated, 3 did not complete the task. Qwen 3.8 from Alibaba got stuck during the process, with no reason stated by Miller. GPT-5.6 Sol from OpenAI and Fable from Anthropic outright refused to work on the code, according to Miller himself, despite the task being to identify and fix flaws, not to exploit them.
The 4 agents that did complete the work, Grok 4.6 (xAI), Kimi K3 (Moonshot AI), DeepSeek V4 Pro, and DeepSeek V4 Flash (DeepSeek), had very different dollar costs for token consumption (user instructions) when executed:
- Grok 4.6 cost $4.70 in computing consumption.
- DeepSeek V4 Pro cost $0.30 for the same task, a difference of 15.7 times (almost 1500%) between the most expensive and the cheapest model to run.
The time taken also varied considerably. DeepSeek V4 Pro completed its work in 30 minutes, Kimi K3 in 40 minutes, and Grok 4.6 in 1 hour, while DeepSeek V4 Flash took 3 hours, six times (500%) longer than DeepSeek V4 Pro for the same task.
According to Miller, only Grok 4.6 matched his own assessment of the severity of the 5 vulnerabilities.
The other 3 agents that completed the work (Kimi K3, DeepSeek V4 Pro, and DeepSeek V4 Flash) exaggerated the severity of the flaws found, something that, according to Miller himself, ultimately affected the overall quality of the results delivered by those 3 models.
Miller's test comes amid a wave of attacks that use artificial intelligence to accelerate the search for and exploitation of security flaws in the Bitcoin ecosystem.
In recent weeks, Coldcard, Boltz, ZEUS, BTCPay Server, and LNP2Pbot suffered security incidents linked to flaws that attackers identified and exploited with the help of AI models. None of these cases affected the Bitcoin protocol itself, but rather products and services operating outside the network.
This wave of incidents prompted the Bitcoin Red Team, a group of developers dedicated to finding security flaws in open-source Bitcoin projects, to launch a massive AI-assisted audit of the ecosystem.
The Bitcoin Red Team audit has already covered more than 300 open-source repositories, and from that work emerged the 5 vulnerabilities that Miller used as the basis for his own test on the 7 AI agents, among another 8,000 flaws found by that group of researchers, as reported by CriptoNoticias.
Of the 7 agents that Miller tested, 4 are open-source, Kimi K3, DeepSeek V4 Pro, DeepSeek V4 Flash, and Qwen 3.8. This means that the companies that trained them publicly released their parameters, allowing any user to download and run them on their own machines without relying on the approval of those companies.
The other 3 agents, Grok 4.6, GPT-5.6 Sol, and Fable, are closed, so they can only be used through the platform or the application programming interface (API) of the company that developed them.
This distinction between open-source and closed models is relevant to the security of the Bitcoin ecosystem. By running on the companies' own servers, closed models allow their creators to maintain active security filters, such as those that led GPT-5.6 Sol and Fable to refuse to work on the vulnerabilities in this test.
Open-source models, on the other hand, can be run on a local computer and modified by any user, making it possible to remove those security filters. An attacker could use an unrestricted version of one of these same models to assist real attacks against platforms in the Bitcoin ecosystem, something that no external filter could prevent.
Miller's test exposes that, beyond the advancement of artificial intelligence in detecting vulnerabilities, significant differences persist between the various models when it comes to classifying and correcting them accurately, both in the dollar cost incurred by each and in the criteria they apply.
-- Price
This content is provided for general informational purposes only and doesn't constitute financial, investment, legal, or tax advice. Any events, rewards, online promotions, or related information mentioned herein should not be considered a recommendation, solicitation, or invitation to purchase, sell, trade, or otherwise deal in any crypto assets. Crypto assets are highly volatile and may result in loss. The availability of WEEX services, products, and related events may vary by region. You are responsible for ensuring that your participation is in accordance with applicable local laws and regulations.
You may also like

U.S. Treasuries, AI, and Inflation: Which Side Will BTC Bet On?

Park Hyun-joo Criticizes Bitcoin Investment Methods

Strategy reduces net leverage to near zero with 6.69 billion dollars in liquidity

5 Bitcoin Supply Signals Behind BTC's Move Toward $80K

Bernstein Predicts Bitcoin Will Reach $150,000 by Mid-2027 and Peak at $300,000 in 2029

Long-term Bitcoin holders increase sales as BTC approaches 80000

Bitcoin: The 'Financial Repression', a New Keyword Driving BTC Upward

Bank Leumi to Launch Cryptocurrency Trading in Early 2027

Bitcoin vulnerable in low-$70,000 range as options market shifts

Bitcoin On-Chain Fund Flow Turns Positive for the First Time, Demand Strength Remains Low

HPC Proposes to SEC and CFTC: Call for Unification of Regulatory Framework for Stock Futures and Perpetual Bond Transactions

Critical Day for Bitcoin: Markets Focused on Data Deluge from the US

Internet Celebrity Di Shi Falls into Trap Set by Acquaintance, Losing Over 10 Million

BitMart Continues Operations Despite End of Trading Service Announcement

Gyndore Public Testnet Launched, Supporting cbBTC Liquidity Center Features

Crypto Fear and Greed Index Reaches 74, Highest Since October 2025

Event Update | Bitcoin Asia 2026 to be Held in Hong Kong from August 27 to 28

Zerohash files second OCC trust bank application

Bitcoin: 8 out of 10 Signals Turn Green in Just One Week

WEEX Exclusive:Bitcoin Peaks Above $81,000 Then Pulls Back | WEEX TradFi Daily (Aug. 26, 2026)

16 Years Ago, 'Stone Man' Lost 9000 BTC in a Legendary Incident! Discover the Lesson Learned

Owning One Bitcoin Is 70 Times Rarer Than Being a Millionaire

Bitcoin +25%, Gold at 3-Month High: Why They're Suddenly Buying the Same Fear
Bitcoin broke $80,000 on August 25. Gold hit a three-month high the same week. Two assets that almost never move for the same reason are suddenly telling the same story — and it's called the debasement trade.

Czech Bitcoin Donation Case Defendant Signals Cooperation, Prosecutors Seek 20-Year Imprisonment

Bitcoin Peaks Above $81,000 Then Pulls Back | WEEX TradFi Daily (Aug. 26, 2026)
Global markets focused on a strong rebound in crypto assets and upcoming U.S. technology earnings. Bitcoin peaked above $81,000 before consolidating at elevated levels, while Ethereum also advanced. Chinese meme tokens and the Layer-2 sector remained active. Technology and semiconductor stocks rebounded ahead of NVIDIA’s earnings, with investors assessing AI infrastructure demand, customer capital expenditure, and earnings conversion. Moderna’s vaccine progress and the decline in crude oil also created significant volatility across related sectors.

Bitcoin payments fade at El Salvador’s Bitcoin Beach

Expansion of US Money Supply Forecasts Long-Term Bitcoin Rise

Mizuho: Current Crypto Rally Quality Surpasses Previous Ones, Driven by Spot and ETFs

Kevin Warsh at Jackson Hole: Why His First Fed Speech Matters So Much for Bitcoin








