The team developing Atria Dawn Preview used AI in 713 of 739 tasks for which participants gave a clear answer about AI use. In the team's study of its own development work, people still made the final choice in 85.5% of recorded decisions about methods or parameters. The paper's title invokes “agentic superintelligence,” but its task records tell a more specific story about how the team divided work.
The preprint was submitted September 14 and revised September 17, 2026. It draws on 769 task records from 56 participants, plus agent logs, in one model development project. The authors describe the records as a case study of how their team used agents to build Atria Dawn Preview.
The big change
- What changed: The team recorded how agents and people shared specific research decisions during a model development project, separating a proposal from the final choice.
- Why it matters: An agent can generate a method and carry out a revision while a researcher still chooses the goal, approves the method and decides whether the result meets the task's criteria. Counting agent actions alone misses those different roles.
- What to watch: These are observations from one team's work, with much of the decision evidence supplied by participants. The findings do not establish how other AI labs divide decisions or whether agents can develop their successors autonomously.
Agents supplied many methods; people selected most
The clearest distinction comes from 567 recorded decisions about methods or parameters among participants in execution roles. The paper reports that AI proposed an option in 64.6% of those decisions. The largest single category, 55.4%, was “AI proposes, human selects.” People made the final choice in 85.5%; AI made it in 9.2%. The remaining decisions were attributed to an external constraint.
The allocation differed by question. People made the final choice in 93.4% of 603 decisions about goals or scope, and 81.9% of 542 decisions about acceptance criteria. AI proposals were far more common for methods than for goals. The paper's figure also records that people selected the final goal in 144 of 151 tasks that participants rated infeasible without AI.
Those percentages describe reported decision roles. They do not measure whether each human choice was well informed. The authors interpret the pattern as researchers steering work while agents generated options within directions people had set. The case study documents delegated execution alongside retained final authority in this project.
Human feedback often sent agents back to work
Participants recalled the most consequential difficulty for each task. Among 588 tasks with a recorded difficulty and response, 447, or 76.0%, moved forward with human help; agents recovered independently in 135, or 23.0%. The main forms of help were extra context or clarified requirements (207 tasks) and a diagnosis or change of method (204). A person took over the main work in four tasks.
The output records show a similar division. Of 627 tasks with a known disposition for the main AI output, 354 required substantive revision. In that revised subset, the agent made the changes after human feedback in 75.4% of cases, while a person edited directly in 19.2%. Researchers could direct a change without making the edit themselves.
One log measure also rose during development: for a cohort of 22 participants, the daily median number of logged agent actions per human prompt went from 11.0 on August 7 to 28.5 on September 4. The measure uses a preceding seven-day window and had valid ratios for 21 or 22 people on each day. It counts activity, so it cannot establish that agents gained authority or that their output improved.
What the task records can establish
Participants estimated whether they could have completed their own part of a task without AI at the same scope, quality and other resources. They rated 151 of 455 completed AI-assisted tasks with usable answers, or 33.2%, infeasible without AI. That is a self-reported counterfactual, not an experiment in which the team repeated those tasks without an agent. It cannot establish that none of the work would have been attempted under different conditions.
The paper does not give a sampling procedure for the 56 participants or the 769 task records, and its public links do not provide the underlying task responses or agent logs for an independent reanalysis. Its benchmark results evaluate Atria Dawn Preview as a model. They do not show an AI system autonomously improving its successor. For teams deciding what to delegate, the reported boundary is specific: agents frequently proposed and revised work in this project, while participants reported retaining most choices about goals, methods and acceptance.



