Bay Area Frontier Research Club #25 | Recursive Self-Improvement (dinner + paper discussion @ The Residency)
About This Event
Recursive self-improvement as an empirical research program.
The claim under examination is narrow: systems that improve the process by which they themselves improve — not merely agents that retry within a task, and not merely human-supervised iteration at higher throughput. That distinction matters, because most of what is currently labeled RSI sits somewhere between those two.
This session treats the problem as a stack with separable research questions: the models and methods that generate improvement, the feedback and evaluation that verify the improvement is real, and the infrastructure that lets the loop run without a human in the middle.
Talks stay short. Discussion is the point. Papers and materials are circulated in advance. Speakers are researchers operating inside these loops; the room is expected to press on assumptions, measurement, and what would count as decisive evidence.
Co-hosted with Inventors Residency of San Francisco.
The Frontier Research Club is a curated forum for rigorous, technical discussion at the frontier of AI. We convene researchers from the frontier labs, Stanford, Berkeley, and the teams building in production to examine concrete work — papers, methods, and results — with a bias toward assumptions, evaluation methodology, failure modes and convincing evidence.
Presentations are intentionally brief so the majority of time is reserved for questions and critique. Materials are shared in advance so the conversation starts at depth.

Agenda
5:30pm: Doors open
5:30pm – 6:30pm: Networking + light dinner
6:30pm – 8:00pm: Research presentations + discussion
8:00pm – 8:30pm: Networking
Presenters & topics
Talk 1: FutureSim — Replaying World Events to Evaluate Adaptive Agents
Shashwat Goel — PhD Researcher, ELLIS / Max Planck Institute Tübingen · Best Paper, ICML Forecasting Workshop · AAAI Outstanding Paper Award
Shashwat works on the science of evaluations and how AI can iteratively improve — the exact hinge of this session. His PhD at Max Planck, advised by Jonas Geiping and co-mentored by Douwe Kiela (CEO, Contextual AI), focuses on designing evaluations and methods to scale AI supervision. Before Tübingen he contributed to Representation Engineering and to WMDP, now the standard benchmark for unlearning dual-use knowledge, and his earlier work earned an AAAI Outstanding Paper Award — top 3 of more than 12,000 submissions.
If an agent is improving itself, how would we know — and can today's frontier agents actually update their beliefs as the world changes?
FutureSim tests that question directly. The environment replays real-world events in chronological order past the model's knowledge cutoff: agents receive daily news over a three-month simulation, maintain a portfolio of forecasts, and decide entirely on their own when to update which beliefs — long-horizon and open-ended, yet reproducible and grounded in real event data. The results separate frontier agents starkly: the best reaches 25% accuracy, several score worse than making no prediction at all, and models with stronger priors start ahead but barely improve as evidence accumulates — scale buys better starting knowledge, not better updating. Shashwat will present the benchmark's design, what the gaps reveal, and the research it makes answerable: test-time adaptation, epistemic humility, memory, search, inference-scaling, and multi-agent self-play.
Pre-read: Goel et al., FutureSim: Replaying World Events to Evaluate Adaptive Agents, Best Paper, ICML Forecasting Workshop.
Talk 2: TBA
Lightning talks — The Residency RSI cohort
Three short talks (3 minutes each) on work in progress from researchers living and building at The Inventors Residency.
Want to present your work?
If you have a research paper you’d like to discuss at one of our next sessions, please submit it for consideration.
Submit your paper here
Who should attend
Researchers working on self-improvement methods, automated research, and evaluation of autonomous loops
Founders and engineers building AI agents and supporting infrastructure
Teams developing autonomous research and model-training workflows
Investors focused on AI infrastructure and research-driven companies
Capacity is limited.

We will take photos and short video clips for event recap and promotion. By attending, you consent to being photographed and recorded, and to the use of those images and clips by the organizers on social media and other event marketing channels.
🌐 Connect with Frontier Research Club
Luma Calendar: luma.com/frontiersyndicate
YouTube: youtube.com/@FrontierResearchClub
Instagram: @frontierresearchclub
Email: kristopher@frontiersyndicate.vc

Hosted by
Frontier Syndicate is a venture community connecting frontier tech researchers, builders, and investors through curated convenings and early-stage capital. Across the Bay Area, we host a recurring series of research forums, builder nights, and intimate investor dinners — and back exceptional companies emerging from the labs, communities, and technical networks we convene.
The Residency is a network of live-in homes for researchers and founders working full-time on frontier problems, with locations across four continents. This session is hosted at The Inventors Residency of San Francisco, its house dedicated to advanced research, home to several researchers working on recursive self-improvement.
SCALE is an equity-free accelerator from GMI Cloud for AI-native startups from pre-seed to Series A. Over six weeks, founders building multimodal, agentic and physical AI get inference credits on NVIDIA-powered infrastructure, hands-on technical support, perks from partners like OpenAI, Notion and Exa, and a Demo Day in front of investors.
Hosted by
We email you a private link showing how many people find this event through Mimetic. Tick the box if you are open to sponsors for events like this one and we will get in touch when there is interest. Anything from there is yours to decide.
Location
📍 San Francisco, CA
Want the other 2 SF Bio events this week? One email, Monday morning.
Unsubscribe anytime.
Tell us who you are and what you sell. A person at Mimetic reads it and gets back to you by email. Nothing is charged here and we do not sell attendee data.
Get a free growth analysis for your company
See how your website, messaging, and go-to-market strategy stack up, in minutes.
Get My Free AnalysisMore SF events like this
AI for Frontier & Meta Science Workshop - Event Description Only, Apply On Website
Kevin Hartz: What Great Founders See Before Everyone Else
AI Insiders: Healthtech Dinner
AI Insiders: AI × Longevity