The most important point in Jerry Tworek's new interview is not his estimate that human researchers may stop being a meaningful part of AI research in roughly two years.

Two years is a bet, not a measurement. The contradiction in his own account is more revealing.

Today's working arrangement is clear: a human formulates a strong hypothesis and chooses the direction; agents write code, run experiments, and collect results. A cycle that once took a month may shrink to a day.

Machine execution of research is becoming fast, cheap, and massive.

Yet Tworek also acknowledges that agents generate many diverse ideas, but their insights are usually shallow. They still do not reproduce the understanding of a person who has spent ten years inside one research field.

His account of the emergence of reasoning models supports this point.

The breakthrough followed years of conviction in scaling reinforcement learning, failed attempts, weak signs of life, a search for suitable data and environments, competition for compute, and a decision to continue before convincing results existed.

That was not merely a series of experiments. It was a continuous research line that preserved intent through uncertainty.

For now, the actual arrangement remains close to a+b: the human preserves direction and recognizes meaningful signals, while machine procedures multiply execution.

But acceleration creates a new bottleneck.

Imagine 100 agents changing the architecture, preprocessing, CUDA kernels, and evaluator at once. The final metric improves by 4%. What improved: the model, the data, the measuring instrument, or a random seed?

If every agent used the same dataset and evaluator, these are not 100 independent confirmations. It may be one shared error counted 100 times as new evidence.

In cybernetic terms, greater variety of machine action requires greater variety of control. Otherwise the system produces convincing traces of error faster than knowledge.

Continual learning adds another question. A system's ability to change during operation does not prove its ability to preserve its organization through those changes.

Which version ran the experiment?
Which line continued after an update?
What happened during a fork, rollback, or migration?
Who owns the accumulated experience?

This is why my work on Temporal AI Presence separates memory, learning, and continuity.

An archive is not continuation.
A restart is not resumption.
Access is not identity.

Tworek is probably right about the direction and may be wrong about the calendar. AI may make experiments cheap before it makes research judgment cheap.

The production of experience will be automated. Its provenance, custody, and continuity will become more important, not less.

The faster research becomes, the more expensive an unnoticed break will be.

Video: youtube.com/watch?si=n63CJSIWik3GSuD_&v=FJfEq9jhpX8&feature=youtu.be