Kaggle agents solved 400 puzzles, winning HSE team a gold medal

Researchers from HSE International Laboratory of Statistical and Computational Genomics and colleagues secured a gold medal at the 2026 NeuroGolf international machine learning championship by deploying an automated system of 330 ChatGPT-based agents. The team overcame a challenge involving 400 unique tasks, each demanding a distinct neural network, making manual completion impractical and necessitating a new approach to machine learning research, Kaggle says.

To manage this system of AI, the researchers also developed a Telegram bot for monitoring and control. “It is no longer a secret that modern models are extremely powerful,” said Aleksei Shmelev, noting that even these powerful models sometimes require encouragement, “like people who are tired after a long working day.”

ChatGPT Agents Solve ARC-AGI Puzzles on Kaggle

The competition, hosted on the Kaggle platform, centered around the ARC-AGI benchmark, a demanding test of abstract reasoning capabilities in AI systems. Participants faced a task: solving 400 unique puzzles, each requiring a distinct neural network architecture tailored to a specific hidden transformation rule within image pairs.

The core of their strategy involved constructing a “factory” of AI agents, each powered by ChatGPT, to work concurrently on a large volume of tasks. The team didn’t simply unleash these agents; they developed a Telegram bot to provide detailed monitoring and control over the entire system, facilitating efficient resource allocation and allowing researchers to prioritize tasks based on their potential for score improvement.

The team tracked the distribution of scores across all 400 tasks, identifying those that yielded the most significant gains when optimized, and then directed the agents’ efforts accordingly, according to Kaggle. This dynamic prioritization was essential given the daily limits imposed on submissions to the Kaggle competition, ensuring each attempt maximized overall progress. Aleksei Shmelev noted that observing the agents during extended periods of autonomous work revealed unexpected behaviors.

“For example, a short prompt such as ‘Don’t give up. You will definitely find a better solution. Think outside the box, explore new branches, and keep trying’ was enough to get the model out of a dead end and eventually find a solution.” The team also found that competitive encouragement could unlock better performance.

Shmelev continued, “After that, the model would quite often immediately find a more optimal solution.” This observation suggests that even advanced AI models can benefit from carefully crafted prompts that address motivational and psychological factors, mirroring human cognitive processes. The researchers published a detailed write-up of their solution on Kaggle, including the architecture of the agent system, implementation details, and the open-source code, making their approach accessible to the wider AI community.

Nikita Chervov emphasized the broader implications of their work, stating, “We have improved our skills in automating work with agents, which we plan to use in the future to speed up the testing of ideas related to our research.” The team intends to leverage this automated pipeline to accelerate their future research endeavors, particularly given the rapidly expanding number of new AI methods and publications, and this approach promises to significantly reduce the time required to formulate, test, and refine research hypotheses, potentially unlocking new breakthroughs in the field.

Stay current

See today’s quantum computing news on Quantum Zeitgeist for the latest breakthroughs in qubits, hardware, algorithms, and industry deals.

Avatar of The Neuron

The Neuron

With a keen intuition for emerging technologies, The Neuron brings over 5 years of deep expertise to the AI conversation. Coming from roots in software engineering, they've witnessed firsthand the transformation from traditional computing paradigms to today's ML-powered landscape. Their hands-on experience implementing neural networks and deep learning systems for Fortune 500 companies has provided unique insights that few tech writers possess. From developing recommendation engines that drive billions in revenue to optimizing computer vision systems for manufacturing giants, The Neuron doesn't just write about machine learning—they've shaped its real-world applications across industries. Having built real systems that are used across the globe by millions of users, that deep technological bases helps me write about the technologies of the future and current. Whether that is AI or Quantum Computing.

Latest Posts by The Neuron: