← Back to Blog
AI NewsBriefing

OpenAI's Unreleased Astra Model Cracks 10 Open Math Problems as Alibaba's Qwen3.8-Max Goes Wide — August 3, 2026

August 3, 2026·9 min read

⚡ Top Story

OpenAI's Unreleased "Astra" Model Solves Ten Long-Standing Open Problems in Math and Theoretical CS — With Machine-Checkable Proofs

On August 1–2, OpenAI disclosed that an internal, not-yet-released version of its next major model, Astra, produced advances on ten open problems spanning mathematics and theoretical computer science — including the first-ever explicit construction of a non-sofic group (a question open since Gromov introduced the concept in 1999), a disproof of Connes's rigidity conjecture, a proof of quantum parallel repetition for general two-player entangled games, and the first improvement to the general sphere-packing exponent since 1978. Unlike prior "AI solves math problem" claims that rest on trusting the lab's word, OpenAI published machine-checkable Lean 4 proof certificates for all ten results on GitHub (Apache 2.0) — independently verifiable by anyone running Lean — reportedly for about $2,000 in compute. Fields Medalist Timothy Gowers reportedly said he'd recommend one of the proofs for a top journal without hesitation.

Why it matters: formal, machine-checked verification is a meaningfully higher evidentiary bar than past AI-math claims, and this is a capability preview from a model OpenAI hasn't shipped yet.

Sources: OpenAI: Ten advances in mathematics and theoretical computer science · Tech Times: OpenAI's Astra solves ten decade-old math problems · The Next Web: OpenAI says Astra has solved ten open problems · Neowin: OpenAI's next major model Astra claims breakthroughs

⚠️ OpenAI's own page could not be directly fetched for verification this session (access blocked); facts cross-validated across four independent outlets citing OpenAI's post and GitHub repo directly. Publish timing (Aug 1 vs. Aug 2) varies slightly by outlet — flagged as borderline against the strict 24-hour window but included given the story's significance and same-week corroboration.


🔬 Research & Papers

No new peer-reviewed academic paper with independently reproducible benchmarks or code was confirmed within the strict 24-hour window — arXiv (cs.AI/cs.LG/cs.CL/cs.CV) and Hugging Face Papers were inaccessible to direct verification this session, so treat this as unverified-empty rather than confirmed-empty. The closest research-adjacent item is OpenAI's Astra math-proof disclosure (see Top Story), though that's a capability announcement with formal proof certificates rather than a traditional paper.


🏢 Industry & Startups

Chip Stocks Post Worst Month Since 2008 Even as Retail Money Pours In

The Philadelphia Semiconductor Index (SOX) fell 21% in July — its worst month since October 2008 — with intraday swings over 2% on every trading session and 4%+ closes on nearly half of them, as investors question whether AI infrastructure capex can keep accelerating at its current pace. Retail investors poured a record $12B into semiconductor ETFs in the most recent week even as institutional sentiment turned cautious. Separately, Alibaba shares jumped roughly 6% on August 3 following the Qwen3.8-Max launch (see Tools & Releases).

Why it matters: it's the first hard signal of systemic investor nervousness about AI-capex sustainability showing up in chip-stock pricing broadly, not just single-company moves.

Sources: Bloomberg: Wall Street's favorite bet comes undone as chips whipsaw market · Fortune: Wall Street's AI trade and the chip-stock selloff · Invezz: Alibaba shares jump 6% after Qwen3.8-Max launch


🛠️ Tools & Releases

Alibaba's Qwen3.8-Max Goes Widely Accessible, Claims It Rivals Anthropic's Fable 5

Alibaba made its flagship Qwen3.8-Max — a 2.4-trillion-parameter mixture-of-experts model (~95B active parameters), multimodal across text/image/video input, with a 1M-token context window — widely accessible to global users on August 3, ahead of a planned open-weights release (including a smaller Qwen3.8-27B) next week. Alibaba claims the model's benchmark scores are "second only to" Anthropic's Fable 5 and that it edges out Moonshot's Kimi K3 on several benchmarks despite using fewer active parameters (95B vs. Kimi K3's 104B); the company also says the model completed a software project independently during a 16-day internal test.

Why it matters: Qwen3.7 to Qwen3.8 shipped in roughly two months — a pace that keeps Alibaba at the front of China's frontier-model race.

⚠️ Alibaba's own benchmark comparisons (vs. Fable 5 and Kimi K3) are the company's claims and have not yet been independently reproduced.

Sources: Bloomberg: Alibaba's Qwen3.8-Max claims benchmark scores rivaling Anthropic · South China Morning Post: Qwen3.8-Max made widely accessible · MarkTechPost: Alibaba Qwen releases Qwen3.8-Max

Thinking Machines Lab Releases Inkling-Small, a Quarter-Sized Open Model That Beats Its Own Teacher

Mira Murati's Thinking Machines Lab released Inkling-Small on August 2 — an open-weights MoE model with 276B total/12B active parameters (about a quarter the size of its July 15 predecessor, Inkling, at 975B total/41B active) that notably surpasses the larger model on some benchmarks: 31.6% vs. 29.7% on Humanity's Last Exam (text-only), and 80.2% vs. 77.6% on SWE-bench Verified. It reasons natively over text, image, and audio with a 1M-token context window; weights ship under Apache 2.0 on Hugging Face.

Why it matters: a smaller model beating its own larger "teacher" on real benchmarks is a notable distillation result, and a fast two-week follow-up intensifies competition in the open-weights space alongside DeepSeek and Kimi.

Sources: MarkTechPost: Thinking Machines Lab releases Inkling-Small · VentureBeat: Thinking Machines debuts Inkling-Small · Hugging Face: Welcome Inkling by Thinking Machines

Microsoft's Project Perception Enters Public Preview Inside Defender

Project Perception — Microsoft's agentic cybersecurity system paired with its in-house MAI-Cyber-1-Flash model (first announced July 27) — entered public preview inside Microsoft Defender on August 3. It coordinates "red" (attack-path mapping), "blue" (threat triage/risk investigation), and "green" (automated remediation) AI agents through an orchestrator; the initial release focuses on software vulnerability management, and autonomous remediation requires human approval for consequential actions.

Why it matters: it's one of the first hyperscaler-shipped agentic security products designed to act on threats, not just alert on them, reaching public preview.

⚠️ Primary Microsoft source pages could not be directly fetched this session; corroborated across multiple independent outlets.

Sources: Windows Forum: Microsoft Project Perception enters public preview August 3 · Futurum: Project Perception bets on agents that act, not just alert


🌏 Global AI & Geopolitics

Alibaba's Qwen3.8-Max wide release (see Tools & Releases) continues China's rapid frontier-model iteration pace — roughly two months between the Qwen3.7 and Qwen3.8 generations. No other China/EU/UK/India/Middle East development was independently verified as genuinely new within the strict 24-hour window.


⚡ Energy, Infrastructure & Chips

See Industry & Startups — the semiconductor sector's worst month since 2008 (SOX -21% in July) is the dominant infrastructure-adjacent story of the window. No new individual chip-supply or data-center power deal was independently confirmed within the last 24 hours.


🤖 AI Agents & Autonomy

Microsoft's Project Perception public preview (see Tools & Releases) — a hyperscaler agentic-security system now live in Defender, with human-gated autonomous remediation.


🔒 Safety, Alignment & Ethics

No new development from NIST/CAISI, the UK AI Security Institute, the Center for AI Safety, the AI Now Institute, Georgetown CSET, Brookings, or the Future of Life Institute was independently verified within the strict 24-hour window.


📊 Numbers & Signals

  • 21% — the SOX semiconductor index's July decline, its worst month since October 2008
  • $12B — record weekly retail inflow into semiconductor ETFs even as the SOX slid
  • 2.4T / ~95B — Qwen3.8-Max's total and active parameters
  • 6% — Alibaba's share-price jump on the Qwen3.8-Max launch
  • $2,000 — reported compute cost for OpenAI Astra's ten math-proof runs, at Sol API rates
  • 276B / 12B — Inkling-Small's total/active parameters, about a quarter of its predecessor Inkling's size
  • 87,714 — AI-linked layoffs tracked so far in 2026, vs. 54,836 for all of 2025 (cumulative tracker figure, not a single-day event — background context only)

🧠 Worth Thinking About

Two stories from today pull in opposite directions. OpenAI quietly demonstrated it can produce independently verifiable mathematical breakthroughs — machine-checked, not just claimed — with a model it hasn't even shipped yet. Meanwhile Wall Street's semiconductor trade, which has been financing the entire AI buildout, just had its worst month since the 2008 financial crisis, even as retail investors kept buying the dip. The frontier of demonstrated capability and the frontier of investor conviction are no longer moving in lockstep, and it isn't yet clear which one is mispriced.


🏛️ Government & Regulation

California's AI Transparency Act Becomes Operative

California's AI Transparency Act (SB 942, as amended by AB 853) took effect August 2, requiring generative-AI providers with 1M+ California monthly users to offer a free public AI-content detection tool and embed machine-readable provenance watermarks in AI-generated images, video, and audio — landing in the same week as the EU AI Act's Article 50 transparency rules.

Why it matters: it's the first major US state-level AI content-transparency regime to actually take effect, not just be signed into law.

⚠️ This is an effective-date story; no enforcement action has been reported yet.

Source: California Legislative Information: SB 942


🔭 Frontier Lab Dispatch

OpenAI — August 1–2: Disclosed that its unreleased Astra model solved ten open math/CS problems with machine-checked Lean proofs (see Top Story).

Microsoft — August 3: Project Perception, its agentic cybersecurity system, entered public preview inside Defender (see Tools & Releases).

No dated August 2–3 post was independently confirmed from Anthropic, Google DeepMind, Meta AI, xAI, Mistral, Cohere, or Hugging Face's own news pages this session.


🔗 Quick Links

Tier 1 — Official / Primary Sources

Tier 3 — Tech & Business Media