Weekly AI News: Science, Robots and Open Models
11 min read

This week AI made news not only with new models but with striking results in science, robotics and open source. A counterexample in plasma physics, the decoding of a 217-year-old coded letter, a humanoid robot playing badminton and several new open models top the list.
This article summarizes the sections of the AI Search channel's weekly AI news video on YouTube that I did not cover in other articles and extends them with project pages and announcement texts; the video link is at the end. This is a news roundup, and I have neither tested nor verified the work. Numbers come from the authors' own pages.
The week's big model and product news is in separate articles: Claude Sonnet 5.5, GPT-6.1 Sol, OpenAI Dots and Gemini 4 Argon.
AI in science: a plasma conjecture and Napoleon's cipher
A counterexample to Grad's conjecture in plasma physics
Fusion reactors have to hold plasma, an extremely hot gas of charged particles, in place with magnetic fields. A reactor type called a stellarator uses very complicated three-dimensional magnetic fields for this. Plasma physics has had an open question for a long time: can plasma sit in a stable, smooth equilibrium inside a 3D magnetic field even when pressure varies from place to place? For a long time it was thought this might only be possible if the field had certain symmetries (Grad's conjecture).
This week two works offered counterexamples to that view. The first is the paper "Counterexamples to Grad's conjecture" by Javier Gómez-Serrano, Lukas Liehr and Mitchell A. Taylor; its GitHub repository contains a formal verification of the main result in Lean 4. The theorem says there is a smooth family of magnetohydrostatic equilibria with only N-fold rotational symmetry; the interior members of the family are not isolated, so it is a continuous family. The second work is Landreman's "analytic 3D MHD equilibria" repository. According to the video, researchers used GPT-6 Astra to find two different unusual configurations in which plasma can stay stable and smooth.
AI's role here is finding counterexamples; producing a result and verifying its proof are separate jobs. Formal verification in Lean means a machine can check the second. The presenter in the video expects such solutions to grow for problems close to pure math: a solution is hard to find and easy to check. That is an opinion. Repositories: Grad-Conjecture, analytic_3d_equilibria.
Napoleon's 217-year-old coded letter
Carter Church published a write-up (September 18, 2026) describing the decoding of a ciphered letter written in 1809 to General Marmont. The letter was known only from a single plate in a 1969 French army journal and was on the list of unsolved historical ciphers. The author's findings:
- Date: Contrary to the wrong date in the journal, the letter is placed in March 1809, in the two weeks before Austria's invasion of Bavaria. On March 16, 1809, Napoleon had ordered his stepson Eugène to send Marmont a ciphered letter.
- Structure: 1,300 cipher units and 155 distinct signs. It is a homophonic cipher: common letters get several signs, so counting letters does not work. Twenty-nine of the signs are whole words, not letters ("de," "que," "les," "vous," "général" and so on).
- Previously known: The table by French cryptology historian Daniel Tant gives only 33 letter values; that covers 435 of the 1,300 units, about a third.
- Process: GPT-6 Astra cut the plate into rows, decided which of 175 hand-drawn marks were the same sign, and then broke the cipher. A solver assigned letters to the remaining signs by simulated annealing and scored each candidate against French letter statistics (3-, 4- and 5-grams). Each sign's proposed value was checked against every occurrence on the plate.
- Time: About six hours of model execution time. All of it came from a single image.
- Verification: The solver was rerun from scratch with Marmont's memoirs and all Napoleonic texts removed from its language model and found the same reading; so the key does not depend on the model having seen that history.
The result is a military briefing laying out the positions of French and allied armies just before Austria went to war. The author also provides the script that regenerates the solution as a downloadable package. Full write-up: Breaking the Marmont Cipher, 1809.
Robotics: a sense of touch and badminton
TactileStep: touch in the feet
Humanoid parkour policies can traverse terrain, but task completion can mask problems such as harsh landings, edge contacts and unstable stance. Humans regulate how softly they touch the ground through pressure feedback in their feet; most humanoid robots lack this sense. TactileStep (CoRL 2026 Spotlight, Tsinghua University) gives the robot a sense of touch through sole pressure sensing. A lightweight tactile simulator turns rigid foot-terrain contact into a pressure array; the policy uses these contact features in both training and on the real robot, and gait phases are estimated online.
Results reported on a Unitree G1:
- Up to a 48.8% reduction in peak touchdown force.
- Up to a 30.1 dB reduction in peak A-weighted impact noise (against a strong perceptive baseline).
- Up to 23.8% more stance contact area.
In the demos the robot walks gently up and down stairs, across slopes, platforms and flat ground. Code is "coming soon." Page: tactilestep.github.io.
A humanoid robot plays badminton
The "Humanoid Badminton" work by researchers from Tsinghua, The Chinese University of Hong Kong and other institutions (accepted at CoRL 2026) trained, according to the video, a Unitree G1 to play badminton. The authors say it is the first real-world humanoid racket-sport system to demonstrate sustained multi-skill human-robot rallies, including highly dynamic jump returns. In the demo the robot does forehands, backhands and jump returns.
The challenge is big: the shuttle moves fast, the hitting window is narrow, the robot must coordinate legs, torso, arm and racket within a few seconds, and usable human motion data is scarce and imperfect. The solution is a three-stage hierarchical reinforcement learning framework: (1) expanding sparsely annotated hitting events into varied stroke variations through task-randomized motion augmentation, forming a continuous skill space, (2) a high-level planner composing skills online according to the shuttle's state, and (3) a context-conditioned adversarial regularizer that encourages natural skill usage. Page: humanoid-badminton.
3D and object interaction: PAMI and Point2Part
PAMI (Tübingen AI Center and the Max Planck Institute for Informatics): From a text prompt and an object, it generates full-body 3D motion of a person interacting with that object. The idea, inspired by the Hough transform, is that body-part anchors "vote" on the object's motion. First PamiGen generates a coarse interaction in a latent space; then PamiRefiner refines the contact geometry using long-range and short-range surface sensors. The aim is to reduce object drift, missed contact and penetration. Tested on the InterAct dataset. According to the video, the code is live, the models are planned and the license is MIT. Page: coral79.github.io/pami.
Point2Part (Carnegie Mellon University): You place one point per desired part, and the model splits the object as a complete partition of the whole. Parts do not overlap and their union covers the entire object. The page's comparison for generating closed parts from a single image:
| Method | Part CD (lower is better) | Part [email protected] | Penetration % | Watertight % |
|---|---|---|---|---|
| Point2Part | 0.0534 | 0.686 | 0.01 | 100 |
| OmniPart | 0.0643 | 0.610 | 0.96 | 94.7 |
| PartPacker | 0.0811 | 0.540 | 1.25 | 86.0 |
| PartCrafter | 0.1275 | 0.300 | 2.45 | 70.5 |
According to the video, the code is live and the model and training code are planned. Page: Point2Part.
New open models
AstaBrief (Ai2)
Ai2 open-sourced AstaBrief 8B on October 2. It is based on Qwen3-8B; it turns a research question and retrieved literature excerpts into a cited report and writes the report in one pass rather than section by section. Training used supervised fine-tuning and direct preference optimization (SFT + DPO) rather than reinforcement learning; the team wanted to try a cheaper recipe than reinforcement learning, which can be unstable and expensive. It is offered as "Fast" mode in Asta's "Generate a report" feature alongside the Claude-powered "Thinking" mode: on average 51.1 seconds per report across the Asta pipeline versus 178.5 seconds in Thinking mode (about 3.5x faster). Weights, training data and an example workflow for making reports from your own PDFs are open. Ai2 also notes that most of the training and evaluation was done in 2025 and the comparison models were that day's frontier models. The full 8B model is about 32 GB; you need a strong graphics card. AstaBrief.
Olmo-core 3 (Ai2)
Not a new chatbot but an open training infrastructure, released October 1, for training large "mixture of experts" (MoE) models. In an MoE only part of the model runs for each input; but the whole model still has to sit in GPU memory, and routing inputs to the right expert creates communication costs across a cluster. Olmo-core 3 switches from the previous fully sharded data parallelism (FSDP) implementation to distributed data parallelism (DDP), keeping experts resident on GPUs and routing the relevant data to them. Measurements on the page:
- On eight NVIDIA B300 GPUs, a 47-billion-parameter MoE processed 52,000 tokens per second per GPU; the earlier implementation managed 19,400 (about 2.7x).
- The expert pool was raised from 8 to 128 while still selecting four experts per token (about 3.2 billion active parameters); total parameters grew from 4.6 billion to 47 billion while training throughput fell by less than 5%.
- The same infrastructure was also benchmarked at over one trillion total parameters.
The technical report, code and an interactive demo are open. Olmo-core 3.
IQuest-Q1
An MoE model for agentic coding, reasoning and multi-step tool use, tuned for use with command-line tools such as Claude Code and Codex. About 320 billion total parameters and about 15 billion active per token; 88 layers, 8 of 256 experts active, 524,288 tokens of context. Text only; not multimodal. The developers recommend the suggested settings (temperature 1.0, top-p 0.95, top-k 20) and using it with Claude Code 2.1.140 or Codex 0.142. For production, SGLang or vLLM with an OpenAI-compatible API. IQuest-Q1.
AREX-2 (BAAI)
The Beijing Academy of Artificial Intelligence's 27-billion-parameter dense long-horizon agent model; Apache 2.0 licensed, 262,144 tokens of context. It learns to improve a solution over several rounds: propose, measure, reflect, revise. It reads scores, logs, errors and timings to decide what to change. It was trained on machine learning and algorithmic programming tasks where the result can be verified; the behavior is said to transfer to deep research too. The model is about 55 GB. AREX 2.
What did we learn?
For an individual developer the most practical item is the open models: IQuest-Q1 and AREX-2 need heavy hardware, while 8-billion-parameter models like AstaBrief are more accessible. The science news shows AI progressing fastest where results are checkable (formal math, code breaking); Lean verification and the "does the key read like French" test are two examples. In robotics the two works follow the same pattern: multiply scarce, imperfect human data in simulation and carry it to a real robot.
Frequently Asked Questions
Did AI disprove the plasma conjecture?
Two works offered counterexamples to that conjecture. According to the video, researchers used GPT-6 Astra to find two different configurations. The result of Gómez-Serrano, Liehr and Taylor was formally verified in Lean 4.
How long did decoding Napoleon's letter take?
About six hours of model execution time, from a single plate image. The letter was a military briefing sent to General Marmont in March 1809.
What does TactileStep provide?
By training with data from pressure-sensing insoles on the robot's feet, it cuts peak touchdown force by up to 48.8% and impact noise by up to 30.1 dB, and raises contact area by up to 23.8%.
Is Olmo-core 3 a model?
No, it is a training infrastructure for training MoE models more efficiently. In Ai2's measurement it gave about 2.7 times the throughput of the previous implementation on a 47-billion-parameter MoE.
Source
- AI Search, Gemini 4, GPT 6.1, Dots, Claude Sonnet 5.5, Ideogram 4.5, Flux 3: AI NEWS (YouTube, October 4, 2026): all sections of this article.
- Carter Church, Breaking the Marmont Cipher, 1809 (September 18, 2026).
- Gómez-Serrano, Liehr and Taylor, Grad-Conjecture (Lean 4 formalization); Landreman, analytic_3d_equilibria.
- TactileStep, Humanoid Badminton, PAMI, Point2Part: project pages.
- Ai2, AstaBrief (October 2, 2026) and Olmo-core 3 (October 1, 2026); IQuest-Q1; BAAI, AREX-2.
Related Posts
What Is GPT-6.1 Sol? Price, Performance, Who Can Use It
GPT-6.1 Sol offers near-Astra intelligence at a fifth of the price: $2 / $10. New plan tiers, Ultrafast speed and who can use it, explained.
Gemini 4 Argon vs Claude Opus 5.5 vs GPT-6 Astra
Gemini 4 Argon leads on legal, finance and long software tasks but trails on terminal and science tasks. Three models compared with a benchmark table.
Whistle and Phonon 2: The Smallest Speech-to-Text Models
Whistle is 16.9 MB and runs on a CPU in seven languages; Phonon 2 transcribes an hour of audio in about 20 seconds on a MacBook Air. Setup and comparison.