Machine Learning Street Talk (MLST) Podcast Por Machine Learning Street Talk (MLST) arte de portada

Machine Learning Street Talk (MLST)

Machine Learning Street Talk (MLST)

De: Machine Learning Street Talk (MLST)
Escúchala gratis

Welcome! We engage in fascinating discussions with pre-eminent figures in the AI field. Our flagship show covers current affairs in AI, cognitive science, neuroscience and philosophy of mind with in-depth analysis. Our approach is unrivalled in terms of scope and rigour – we believe in intellectual diversity in AI, and we touch on all of the main ideas in the field with the hype surgically removed. MLST is run by Tim Scarfe, Ph.D (https://www.linkedin.com/in/ecsquizor/) and features regular appearances from MIT Doctor of Philosophy Keith Duggar (https://www.linkedin.com/in/dr-keith-duggar/).Machine Learning Street Talk (MLST)
Episodios
  • How Researchers Test AI for Hidden Goals — Apollo Research
    Jul 31 2026

    Can an AI do the right thing for the wrong reason? Tim Scarfe speaks with Apollo Research’s Alexander Meinke, Axel Højmark and Jérémy Scheurer about Measuring Reward-Seeking via Contrastive Belief Updates, their new research with OpenAI.


    The panel asks how models infer what graders reward, why good behaviour can come from the wrong reason, and whether that difference can be measured. The conversation moves through promise-breaking, grader awareness, reward hacking, scheming, opaque reasoning and corrigibility, then turns to a detailed walkthrough of the contrastive-belief method and what its results do and do not show. The o3 results discussed here concern an intermediate checkpoint without safety training.


    This episode was made in partnership with Apollo Research. MLST retained full editorial control.


    Reference

    Apollo Research: https://www.apolloresearch.ai/


    ---

    TIMESTAMPS:

    00:00:00 Cold Open

    00:02:12 Right Things, Wrong Reasons

    00:12:47 Grader Awareness

    00:26:22 Legibility

    00:32:35 What To Call It

    00:35:58 Intelligence, Agency, Anthropomorphism

    00:45:16 Apollo’s Mission

    00:48:54 The End of the Exponential

    00:55:45 The Paper

    01:16:34 Closing Reflection


    ---

    REFERENCES:

    tool:

    [00:00:08] Claude Fable

    https://www.anthropic.com/claude/fable

    [00:12:50] AlphaGo Zero

    https://deepmind.google/blog/alphago-zero-starting-from-scratch/

    [00:44:30] AlphaFold 3

    https://deepmind.google/science/alphafold/

    paper:

    [00:01:02] Measuring Reward-Seeking via Contrastive Belief Updates

    https://arxiv.org/abs/2607.18966

    [00:16:19] Natural Language Autoencoders Produce Unsupervised Explanations of LLM Activations

    https://transformer-circuits.pub/2026/nla/

    [00:26:48] Stress Testing Deliberative Alignment for Anti-Scheming Training

    https://arxiv.org/abs/2509.15541

    [00:35:33] Shortcut learning in deep neural networks

    https://arxiv.org/abs/2004.07780

    [00:53:49] Measuring AI Ability to Complete Long Software Tasks

    https://arxiv.org/abs/2503.14499

    [00:59:52] Modifying LLM Beliefs with Synthetic Document Finetuning

    https://alignment.anthropic.com/2025/modifying-beliefs-via-sdf/

    [01:10:44] Alignment Faking in Large Language Models

    https://arxiv.org/abs/2412.14093

    [01:13:55] Natural Emergent Misalignment from Reward Hacking

    https://www.anthropic.com/research/emergent-misalignment-reward-hacking

    other:

    [00:10:14] We Need a Science of Scheming

    https://www.apolloresearch.ai/science/science-of-scheming/

    [00:32:56] CoastRunners reward hacking example

    https://deepmind.google/blog/specification-gaming-the-flip-side-of-ai-ingenuity/

    organization:

    [01:06:07] Redwood Research

    https://www.redwoodresearch.org/


    ---

    ReScript:

    https://app.rescript.info/share/718ab68e18cfa3b9b800da6b3290fd42

    Más Menos
    1 h y 19 m
  • Why a Nation Can't Outsource Its Frontier AI - Alistair Pullen (Cosine AI)
    Jul 13 2026

    This episode is sponsored by Notion. Learn more about Notion's Developer Platform today at https://notion.com/mlst


    Britain's most capable coding model can't be exported, and that ban is the whole reason Cosine set out to build one from scratch. Alistair Pullen, CEO and co-founder of Cosine, sits down with Tim Scarfe to explain how a frontier system he calls Fable, locked behind US export controls, became the founding case for a UK sovereign model trained on the Isambard supercomputer in Bristol.


    The bet underneath it is economic. Pullen argues that an inference company, rather than a training-first lab, doesn't need billions to compete: millions, a national compute allocation, and a consortium feedback loop can be enough. From there it gets into the machinery, why open-weight models still trail the frontier on size, active parameters and data, the mixture-of-experts versus dense trade-off and why active params dominate how a model actually feels, and the edge that real coding trajectories confer.


    The back half is about making agents trustworthy. Pullen makes the case for beating "slop" by rewarding the process instead of the final answer, reframes code review as runtime proof (spin the bug up in a VM and force the agent to actually exploit it), and walks through Swarm, Cosine's system running hundreds of sub-agents in one shot. It ends on why memory is still an unsolved hack, how synthetic graders let you run RL on tasks with no built-in test, and why Pullen reads US export controls as an accidental gift, with a supply-chain sting in the tail.


    ---

    TIMESTAMPS:

    00:00:00 The sovereign mandate and the Fable ban

    00:04:02 Millions vs billions: the inference-company model

    00:07:19 The consortium feedback loop

    00:07:40 Why open models lag the frontier

    00:14:59 MoE vs dense, and why active params matter

    00:16:29 Trajectories: the process-data advantage

    00:19:48 Beating slop: reward the process, not the answer

    00:26:06 Reusable abstractions and the epistemic wall

    00:29:56 Code review becomes runtime proof

    00:37:32 Do agentic harnesses still matter?

    00:40:35 Swarm: orchestrating hundreds of sub-agents

    00:45:14 Why memory is still unsolved

    00:48:25 Synthetic data and graders for RL

    00:53:09 The US export gift and supply-chain risk


    ---

    REFERENCES:

    organization:

    [00:01:15] Cosine

    https://cosine.sh

    [00:04:14] Mistral AI

    https://mistral.ai

    [00:05:50] Anthropic

    https://www.anthropic.com

    [00:07:42] Cohere

    https://cohere.com

    [00:08:36] DeepSeek

    https://www.deepseek.com

    tool:

    [00:02:52] Isambard-AI

    https://isambard.ac.uk

    [00:05:56] Colossus (xAI)

    https://en.wikipedia.org/wiki/Colossus_(supercomputer)

    [00:07:52] GLM (Z.ai)

    https://z.ai

    [00:11:52] NVIDIA B300

    https://www.nvidia.com/en-us/data-center/dgx-b300/

    [00:15:37] gpt-oss-120b

    https://huggingface.co/openai/gpt-oss-120b

    [00:15:52] Devstral 2

    https://mistral.ai/news/devstral

    [00:16:01] Llama 70b

    https://www.llama.com

    [00:17:05] Claude Code

    https://www.anthropic.com/claude-code

    [00:26:23] ARC-AGI (Francois Chollet)

    https://arcprize.org

    [00:40:38] Swarm (Cosine)

    https://cosine.sh

    [00:40:50] OpenAI Codex

    https://github.com/openai/codex

    [00:41:16] Lumen Outpost (Cosine)

    https://cosine.sh

    [00:41:18] Kimi K2 (Moonshot)

    https://huggingface.co/moonshotai/Kimi-K2-Instruct

    [00:49:55] SWE-bench

    https://www.swebench.com

    [00:52:40] SystemVerilog

    https://en.wikipedia.org/wiki/SystemVerilog

    person:

    [00:23:40] Andrej Karpathy

    https://karpathy.ai

    paper:

    [00:27:10] GRPO (DeepSeekMath)

    https://arxiv.org/abs/2402.03300

    [00:27:13] GSPO

    https://arxiv.org/abs/2507.18071


    Incompressible Knowledge Probes, Bojie Li

    https://arxiv.org/pdf/2604.24827


    Estimating the Size of Claude Opus 4.5/4.6

    https://unexcitedneurons.substack.com/p/estimating-the-size-of-claude-opus


    ---

    ReScript:

    https://app.rescript.info/session/5852d2b884c4ce4b?share=10b9799160845bb11779f8ac6cd3124f

    Más Menos
    56 m
  • The Benchmark With No Instructions — ARC-AGI-3 (winning team!)
    Jul 1 2026

    Tim Scarfe travels to Zurich to sit down with the Tufa Labs ARC-AGI-3 team — founder Benjamin Crouzier, with Jeroen Cottaar, Dries Smit, Stefano Viel and Michal Tesnar — to work out what their leaderboard-topping system does and what the benchmark is really testing.The cut opens on the games: a walkthrough of the Locksmith game, where you read the rules of an unfamiliar world straight from raw frames. ARC-AGI-3 makes ARC interactive and agentic, so the model has to *discover* the goal rather than transduce a static grid. It stays easy for humans and breaks LLMs, and it runs through everything that follows. Dries traces his StochasticGoose preview win — brute force that only searched actions which changed the frame — and why it collapsed once the organisers added action-efficiency scoring and unseen games.Induction and transduction run through the middle of the conversation — how much of an answer is really priors leaking back the moment a model recognises a maze. The abstraction mountain, and Tim's case that LLMs reach the right answer through fractured, entangled representations — performance, not competence. Whether transformers plan at all or just fake it well enough. Why the score really measures action efficiency, not games solved, and why agents lock onto the wrong goal and cannot climb back out.Crouzier closes on the Tufa Labs thesis — a small lab against the giants, the bitter lesson against hand-built harnesses, and safety — and Tim ties it back to Kenneth Stanley, deep constraints, and creativity as competence.


    Disclosure: Tufa Labs sponsors MLST. ---TIMESTAMPS:00:00:00 Meet the Tufa team and what makes ARC-AGI-3 hard00:02:11 Locksmith game: reading the rules from raw frames00:03:10 Why build an independent research lab00:04:11 StochasticGoose: a preview win, then the hardened games00:07:58 Induction, transduction, and priors inside LLMs00:10:31 Curiosity, world models, and exploring by frame change00:14:32 Understanding debt and losing sight of your own code00:15:53 Requirements-based agents and human-AI co-creativity00:19:22 Why auto-research misses the big picture00:21:54 The abstraction mountain and fractured representations00:27:36 Constraints and making LLMs act as if they understand00:34:51 Human difficulty calibration, esports priors, and emergence00:41:35 Agency, goal acquisition, and two kinds of planning00:47:31 Harnesses, the 36% number, and wrong-goal loops00:52:33 Rewards, goals, and why ARC-AGI-3 resists brute force01:00:46 Would solving ARC-AGI-3 prove AGI?01:07:53 Stripping language away, then priors leak back01:14:06 Representation and whether language is necessary01:18:04 The bitter lesson versus specialised harnesses01:22:20 Capability research, safety, and the software singularity---REFERENCES:organization:[00:02:11] ARC-AGI-3https://arcprize.org/arc-agi/3[00:03:10] Tufa Labshttps://tufalabs.ai/team/[00:04:20] ARC-AGI-3 Preview Agent Competitionhttps://arcprize.org/competitions/arc-agi-3-preview-agentstool:[00:04:55] StochasticGoose ARC-AGI-3 solutionhttps://github.com/DriesSmit/ARC3-solution[00:07:42] ArcGenticahttps://github.com/symbolica-ai/arcgentica[00:07:49] RGB-Agenthttps://github.com/alexisfox7/RGB-Agent[00:14:38] Claude Codehttps://www.anthropic.com/claude-code[01:03:42] Qwen 3.6 27Bhttps://huggingface.co/Qwen/Qwen3.6-27Bpaper:[00:13:03] On the Measure of Intelligencehttps://arxiv.org/abs/1911.01547[00:27:42] DreamCoderhttps://arxiv.org/abs/2006.08381[00:43:55] On the Biology of a Large Language Modelhttps://transformer-circuits.pub/2025/attribution-graphs/biology.html[01:18:46] ImageNet Classification with Deep CNNs (AlexNet)https://proceedings.neurips.cc/paper_files/paper/2012/file/c399862d3b9d6b76c8436e924a68c45b-Paper.pdfother:[01:18:16] The Bitter Lessonhttp://www.incompleteideas.net/IncIdeas/BitterLesson.html---https://app.rescript.info/share/463d7f031349b4b9db428553eed88230

    Más Menos
    1 h y 25 m
adbl_web_anon_alc_button_suppression_t1
Todavía no hay opiniones