
@_akhaliq
AI research paper tweets, ML @Gradio (acq. by @HuggingFace 🤗) dm for promo ,submit papers here: https://t.co/UzmYN5XOCi
RRSI Regularized Recursive Self-Improvement of Agent Harnesses paper: huggingface.co/papers/2609.24…
WorldCrafter Consistent Video World Model with Implicit 3D-aware Memory paper: huggingface.co/papers/2609.24…
CodeMidas Scaling Agentic Coding RL Environments from Code Itself paper: huggingface.co/papers/2609.22…
Qwen-Image-2.1 is out on Hugging Face app: huggingface.co/spaces/akhaliq…
JEPA-Anything Learning Predictive Models across Different Worlds paper: huggingface.co/papers/2609.20…
An Empirical Study of Harness Design for Coding Agents paper: huggingface.co/papers/2609.20…
Agora Git as Shared Memory for Collective AutoResearch paper: huggingface.co/papers/2609.18…
LimiX-2 A Contextual Mechanism Network Towards General Structured-Data Intelligence paper: huggingface.co/papers/2609.17…
StepAudio 3 Realtime Technical Report paper: huggingface.co/papers/2609.14…
Continual Learning Mechanisms Compose for Long-Horizon Memorization paper: huggingface.co/papers/2609.06…
Vidu S2 Real-Time Interactive, Editable, and Spatial Video Generation paper: huggingface.co/papers/2609.11…
SAS Simple Attention Sparsification via End-to-End Optimization of Context Ranking paper: huggingface.co/papers/2609.13…
AuK Technical Report An Open-Source Foundational Model for Speech Generation and Editing paper: huggingface.co/papers/2609.08…
Marigold V2 Revisiting Diffusion Transformers for Monocular Depth Estimation paper: huggingface.co/papers/2609.08…
NeoHorse-1 Towards Recursive Self-Improvement via Agentic Post-Training with Routing Harness paper: huggingface.co/papers/2609.08…
Unlocking Lossless Speedups in LLMs via Discrete Diffusion paper: huggingface.co/papers/2609.04…
It Takes Two to Match Co-Evolving Generative Retriever with Reinforcement Learning paper: huggingface.co/papers/2609.00…
HarnessDev Can LLMs Create and Evolve Their Own Agent Harness? paper: huggingface.co/papers/2609.01…
Repo-To-Skill Distilling GitHub Repositories Into AI4AI Skills paper: huggingface.co/papers/2609.02…
Does On-Policy Distillation Really Distill? From Noisy Teacher to Self-Improvement paper: huggingface.co/papers/2608.31…
LoopArena Benchmarking Models as Runtime Controllers for Loop Engineering paper: huggingface.co/papers/2608.28…
Code as Worlds Agentic Discovery of Executable World Representations for Physical Reasoning paper: huggingface.co/papers/2608.27…
VGI-Bench Probing Visual Intelligence in Video Generation Models paper: huggingface.co/papers/2608.19…
Autonomous Mathematical Discovery in an Open-World Multi-Agent Environment paper: huggingface.co/papers/2608.23…
Apodex 1.1 Scaling Agentic Intelligence for Complex Work paper: huggingface.co/papers/2608.23…
Learning How the World Evolves Extrapolative Video World Models via Latent Dynamics Reasoning paper: huggingface.co/papers/2608.09…
InfinityEdit Infinite Video Editing with a Lightweight Edit-Ignition Adapter paper: huggingface.co/papers/2608.20…
MiniMax-Music3 gradio workflow is out huggingface.co/spaces/akhaliq…
SCoPE Sightline-Coordinate Positional Encoding for Video Diffusion Transformers model: huggingface.co/TencentARC/SCo…
BDH-CQ In-Context Learning with Recurrent Latent Reasoning paper: huggingface.co/papers/2608.09…
SWE-Bench ProMax Benchmarking Agents on Large-Scale Multilingual Code Refactoring paper: huggingface.co/papers/2608.09…
MatrAIx Simulating the World with 8.3 Billion Persona Agents paper: huggingface.co/papers/2608.04…
Toward Skill-Native LLMs Skill Entropy for Benchmarking and Training Long-Horizon Reasoning paper: huggingface.co/papers/2608.05…
MiniMax-H3-Turbo-Lora huggingface.co/spaces/akhaliq…
without the lora
Towards Physics of Multimodal Pretraining Knowledge Flow, Modality Synergy, Early Unification, and Recipes paper: huggingface.co/papers/2608.05…
MerchantBench Benchmarking LLM Agents for Long-Term Coherence in E-Commerce Operations paper: huggingface.co/papers/2607.28…
DAPD Dual-Anchored Policy Distillation paper: huggingface.co/papers/2608.01…
LongHorizon-Harness Advancing Long-Horizon Agents for Real-World Tasks paper: huggingface.co/papers/2608.01…
From RLVR to RLSVR Task Transformation Induces Self-Verifiable Rewards for Open-Ended LLM Self-Improvement paper: huggingface.co/papers/2607.23…
Qwen-UI-Agent Technical Report Toward Next-Generation Real-World Centric Foundation GUI Agents paper: huggingface.co/papers/2607.28…
Explorative Modeling Unlocking a Third Pretraining Axis and End-to-End Generation paper: huggingface.co/papers/2607.27…
DeepSeek-V4-Flash-0731 is out huggingface.co/deepseek-ai/De…
ID-V2V Identity-preserving Video Restylization paper: huggingface.co/papers/2607.22…
CoRT Counterfactual Replay for Token-Level Rubric-Guided Policy Optimization paper: huggingface.co/papers/2607.25…
TurboVLA Real-Time Vision-Language-Action Model at 32 Hz on an RTX 4090 with <1 GB VRAM paper: huggingface.co/papers/2607.27…
A New Role for Relevance Guiding Corpus Interaction in Agentic Search paper: huggingface.co/papers/2607.24…
HiFi-UMI Learning Deployable Manipulation Policies from High-Fidelity UMI Data Alone paper: huggingface.co/papers/2607.25…
A.X K2 just dropped on Hugging Face Large-Scale Sparse MoE (688B / 33B Active) huggingface.co/skt/A.X-K2
From Proprietary to Open-Source Bridging the Distribution Gap via Multi-Agent Protocol Distillation in Agentic Search paper: huggingface.co/papers/2607.24…
StateAct Program State, before Pixels, for Long-Horizon Computer-Use Agents paper: huggingface.co/papers/2607.22…
kimi k3 in claude code via hf claude
Kimi K3 is out huggingface.co/moonshotai/Kim…
Apple-π Benchmarking Thinking with Video Towards Law-Grounded Physical Intelligence paper: huggingface.co/papers/2607.16…
SLAI T-Rex Full-Parameter Post-training of the DeepSeek-V4 Family on Ascend SuperPOD paper: huggingface.co/papers/2607.20…
Reading and Steering Representations of Materials Science Mechanisms in an Open Weight Language Model paper: huggingface.co/papers/2607.20…
DataFlow-Harness A Grounded Code-Agent Platform for Constructing Editable LLM Data Pipelines
Solar Open2 250B just dropped on Hugging Face huggingface.co/upstage/Solar-…
ABot-World-0 Infinite Interactive World Rollout on a Single Desktop GPU
Apple presents Environment-free Synthetic Data Generation for API-Calling Agents huggingface.co/papers/2607.16…
Motif-3-Beta just dropped on Hugging Face ~314B total parameters / ~13B active per token (sparse MoE) 256K context length (262,144 tokens), natively long-context Sparse routing: 384 experts with 8 activated per token, plus 1 shared expert Multilingual, general-purpose t.co/EqRUYr4OGD
RESOURCE2SKILL Distilling Executable Agent Skills from Human-Created Multimodal Resources
VideoChat3 Fully Open Video MLLM for Efficient and Generalist Video Understanding
thinkingmachines Inkling is now available in claude code via hf claude
Harness Handbook Making Evolving Agent Harnesses Readable,Navigable, and Editable
Read It Back Pretrained MLLMs Are Zero-Shot Reward Models for Text-to-Image Generation
Weak-to-Strong Generalization via Direct On-Policy Distillation
Bonsai 27B now available in claude code via hf claude
prism-ml/Ternary-Bonsai-27B now available in claude code via hf claude with @togethercompute
Scalable Visual Pretraining for Language Intelligence
Long-Horizon-Terminal-Bench Testing the Limits of Agents on Long-Horizon Terminal Tasks with Dense Reward-Based Grading
Video Generation Models are General-Purpose Vision Learners
OPSD-V On-Policy Self-Distillation for Post-Training Few-Step Autoregressive Video Generators
Vidu S1 A Real-Time Interactive Video Generation Model
LingBot-World 2.0 (Infinity) is out on Hugging Face interactive world model with: Hour-long generation with zero quality drift Rich actions & events: attack, cast spells, shoot, summon storms Agentic world: a Director Agent drives real-time world evolution 720p/60fps. Playable like a game
LingBot-Video is out on Hugging Face MoE-based video foundation model built for embodied intelligence 30B params, only 3B active at inference Augmented with 70K hours of embodied data on top of large-scale internet video pretraining
LingBot-Video is out on Hugging Face MoE-based video foundation model built for embodied intelligence. 30B params, only 3B active at inference Augmented with 70K hours of embodied data on top of large-scale internet video pretraining
RynnWorld-4D 4D Embodied World Models for Robotic Manipulation
Gemma 4 Technical Report