In an MIT lab, there is a six-qubit quantum chip that spends its time drifting. It comes uncalibrated, and getting it back into shape requires hundreds of measurements, sometimes thousands. A researcher from the group that made it spends several days on it, identifying the right frequencies, adjusting the pulses and checking that the information is still there. Someone has just done that work in their place.
OpenAI described the experiment on September 9, and the lab is a co-signer. The group is called Engineering Quantum Systems, it is hosted at MIT, and the experiment was carried out by Beatriz Yankelevich, a doctoral student. The program that worked on the chip is the same one that writes code at your place: GPT-5.6 Sol, the model behind Codex. The doctoral student gave it the lab's instructions and the chip's plans, then let it take control.
What it does between two measurements
You don't program a quantum computer like a normal computer. The computing unit is called a qubit, it's the equivalent of the bit in our machines, except that it isn't worth zero or one but both at once, and above all that it loses its information in a few millionths of a second if you leave it alone. To make it work, you put it in a giant freezer, a cryostat, at a temperature colder than outer space. And to talk to it, you use microwave pulses, like those from an oven, but infinitely finer.
These pulses have to be adjusted. Every chip is different, every qubit has its own frequency, and the slightest speck of dust on a contact shifts everything. This adjustment is what keeps the researcher busy, and it's exactly what the agent did all by itself: it chooses the parameters for the next measurement, it controls the hardware through the lab's software, it looks at what comes out, and it decides either to start again, or to keep the result and move on. The three stages of the job were finding the transition frequencies, calibrating the pulses, and measuring how long the information lasts before disappearing.
Six qubits, a freezer colder than outer space, and an adjustment that never holds for long
The loop is nothing exotic: it's the one used by any reasonably serious agent. Measure, understand, decide, start again. What changes everything here is what's at the end of the wire. It's no longer a database or a website, it's a machine worth several million sitting in a temperature-controlled room.
What works, and what still doesn't
The nuance matters more than the headline. OpenAI writes it in black and white that the agent does well when the signals are clear, and that it degrades badly when the data is weak, noisy or ambiguous. In other words, on a clean problem, it moves forward on its own. Faced with a curve that makes no sense, it flounders, and a human takes back control.
The company also acknowledges that experienced researchers still find the best settings faster than current models. And the demonstration has its limits: one chip, six qubits, one standard type of machine, and procedures that the lab already knew by heart. The experiment doesn't prove that an agent could diagnose a never-before-seen hardware failure all by itself or interpret an unprecedented result. That's the most important limitation in the whole case and I keep it in mind.
When the signal is clean, it moves forward on its own. When the curve goes all over the place, the researcher takes back control
The real bottleneck in quantum computing
People are always talking about the number of qubits as if it were a speedometer reading. That's a misreading, and I wrote an entire article about it last month: what's holding this technology back isn't making the qubits, it's everything needed around them to keep them going and make them communicate. The cables, the room, the cold, and the tuning. Especially the tuning.
Think about what that means at the scale of a useful machine. A six-qubit chip takes days to fine-tune. A machine with a million of them, the figure we read everywhere for breaking today's encryption codes, would require a million fine adjustments. Nobody is going to do that by hand, not even with an army of PhD students.
That's where this kind of experiment gets its value. Not because an AI would have done new physics, it hasn't done any. Because calibration work is exactly the kind of repetitive, patient task that we don't know how to automate with conventional scripts and that an agent can take on. What was a human obstacle becomes a software problem again.
Concretely, what does it change for you
Nothing this year, and probably nothing in the next ten. Might as well say it right away, because this technology has been dragging around a reputation for permanent miracles for twenty years.
What it promises, though, is very concrete the day it works. Simulating a molecule to find out whether it treats something, without making it or testing it on sick people. Searching for the right combination of materials for a battery without making three thousand prototypes. Optimizing truck routes or power grids. These are calculations that become impossible by getting too big, and that would become doable.
The day the machine delivers on its promises, the first box to change shape might be that one
The honest timeframe, the one nobody can give you: ten years, twenty years, or never. The machines exist, they work, they do things a conventional computer can't do, and they're still too small to be useful for anything other than research. That's exactly where electricity was when people used it to make sparks in a sitting room.
The anecdote that makes me smile is that the first lasers spent ten years being described as a solution in search of a problem. Today, there's one in your supermarket checkout, one in your disc player and a whole forest of them in the cables that bring you the Internet.
What I take away from this story
It's not the AI's performance that strikes me, it's the kind of work it was given. We stopped asking it to talk and asked it to tune a machine. That shift matters more than the lab result.
And there's a question nagging at me. If an agent knows how to tune a quantum chip, how long before another one tunes the boiler, the gate and Mr. Everybody's entire house? I'm asking because I believe the answer won't come down to the model's intelligence. It will come down to the number of people willing to let a program press a button in their place.



Join the conversation
You need an account to comment on this article. Creating one is free and takes under a minute.
No comments yet.