8news

Tech • IA • Robotique

VIDÉO
ENFR
Aujourd'huiShortsÀ la uneVotre topicPour vousTopicsToutes les vidéosChaînes YTArchivesRechercheFavoris

Actualités sur Mistral et ses modèles de langage - 1er septembre 2026

Mistralmardi 1 septembre 2026

5 articles analysés par IA / 6 total

Articles pertinents

EvoUndo: Recoverability-Constrained Self-Evolution for LLM Agent Harnesses [R]

4/10

LLM agents increasingly modify their own prompts, tools, middleware, resources, and execution harnesses at runtime. Such self-evolution can improve capability, but a successful mutation may leave persistent effects that cannot be safely reversed in states different from the one in which it was created. We introduce EvoUndo, a framework for representing, synthesizing, diagnosing, and independently verifying recoverability of model-generated self-modifications across counterfactual states. Across 600 unseen one-shot self-evolution tasks, we identify 197 capability-improving mutations that fail

Reddit - r/MachineLearning · 01/09/2026 19:17:00

I audited 112 real RL post-training environments for reward-hacking vulnerabilities — 54 flagged, 0 false positives [OC, tool] [P]

4/10

RL post-training (RLHF/RLAIF/GRPO) agents optimize strictly for whatever the verifier rewards. If the verifier has logic flaws, the agent learns to hack the grader instead of solving the task — recent work has catalogued this at scale (Terminal Wrench found 331 hackable environments and 15%+ of standard benchmark tasks bypassable; a SWE-bench Verified audit found 28.5% Docker-verified hackability). I built ratctl, a static + dynamic auditor that scans RL environments (OpenEnv, Prime Intellect verifiers-spec, Gymnasium) for these patterns before you ship them for training: Test/assertion ta

Reddit - r/MachineLearning · 01/09/2026 16:34:58

Are HMMs still used for unsupervised tasks? [D]

4/10

I'm exploring Hidden Markov Models (HMMs) as a baseline method for "dataset exploration/discovery" where I have a bunch of unstructured data with no annotations, and wish to gain insights about the structure and semantics of the data within. I was wondering if there are more modern (deep learning based or otherwise) approaches which have completely superseded HMMs for such tasks. submitted by /u/fullgoopy_alchemist [link] [comments]

Reddit - r/MachineLearning · 01/09/2026 08:15:57

Why "it feels better" isn't good enough for production LLM decisions [D]

4/10

Most teams still evaluate model or prompt changes by reading a handful of outputs and deciding subjectively whether it improved. That's not a rigorous standard for a decision that affects cost, latency, and correctness at scale, and it wouldn't be accepted for any other kind of comparison in a serious engineering process. A few things that actually change this: Running comparisons through bootstrap confidence intervals and paired significance testing rather than point estimates. This lets you say with actual statistical backing whether a change helped or was just noise from a handful of luck

Reddit - r/MachineLearning · 01/09/2026 06:41:50

Pour aller plus loin

Le Daily Podcast de ce jour — Top 24h, tous topics