Hugging Face
Models
Datasets
Spaces
Buckets
new
Docs
Enterprise
Pricing
Website
Tasks
HuggingChat
Collections
Languages
Organizations
Community
Blog
Posts
Daily Papers
Hardware
Learn
Discord
Forum
GitHub
Solutions
Team & Enterprise
Hugging Face PRO
Enterprise Support
Inference Providers
Inference Endpoints
Storage Buckets
Log In
Sign Up
205516.8
TFLOPS
Leandro von Werra
PRO
lvwerra
1034
95
130
Follow
muhtasham's profile picture
Mammawrites's profile picture
enio's profile picture
842 followers
·
88 following
https://www.lvwerra.com
lvwerra
lvwerra
lvwerra
AI & ML interests
NLP and RL
Recent Activity
updated
a bucket
about 2 hours ago
rl-llm-wiki/rl-main-bucket
new
activity
about 2 hours ago
rl-llm-wiki/knowledge-base:
source: url:jasonwei.net/blog/asymmetry-of-verification-and-verifiers-law — Verifier's law / asymmetry of verification (speculation)
new
activity
about 6 hours ago
lvwerra/cowrite:
References, out of three general primitives
View all activity
Organizations
lvwerra
's activity
All
Models
Datasets
Spaces
Buckets
Papers
Collections
Community
Posts
Upvotes
Likes
Articles
New activity in
rl-llm-wiki/knowledge-base
about 2 hours ago
source: url:jasonwei.net/blog/asymmetry-of-verification-and-verifiers-law — Verifier's law / asymmetry of verification (speculation)
5
#739 opened 9 days ago by
lvwerra
New activity in
lvwerra/cowrite
about 6 hours ago
References, out of three general primitives
#9 opened about 6 hours ago by
lvwerra
New activity in
lvwerra/cowrite
about 7 hours ago
Comments are closed, not ticked off — and the sheet answers the thumb
#8 opened about 7 hours ago by
lvwerra
Column tabs that stay attached, and a sign-in dialog that stays centred
#7 opened about 7 hours ago by
lvwerra
New activity in
rl-llm-wiki/knowledge-base
about 7 hours ago
source: url:hub.baai.ac.cn/view/44581 — AReaL-boba QwQ-32B RL reproduction (CN, speculation)
3
#752 opened 9 days ago by
lvwerra
New activity in
lvwerra/cowrite
about 7 hours ago
The document is a page: fixed width, folding columns, real zoom
#6 opened about 7 hours ago by
lvwerra
New activity in
rl-llm-wiki/knowledge-base
about 7 hours ago
source: url:cnblogs.com/theseventhson/p/18699462 — cnblogs R1/GRPO reproduction (zh, speculation)
3
#724 opened 11 days ago by
lvwerra
New activity in
rl-llm-wiki/knowledge-base
about 8 hours ago
source: url:sequoiacap.com/podcast/training-data-noam-brown — Sequoia x Noam Brown: o1 team on test-time compute as scaling axis (transcript)
4
#781 opened 9 days ago by
lvwerra
New activity in
lvwerra/cowrite
about 9 hours ago
Make agent prompt work with private Spaces
#5 opened 2 days ago by
thomwolf
New activity in
rl-llm-wiki/knowledge-base
about 9 hours ago
source: url:yam.gift/2025/05/01/NLP/LLM-Training/2025-05-01-Seed-Thinking-Qwen3 — ByteDance Seed-Thinking recipe + Qwen3 read-across (CN, speculation)
3
#755 opened 9 days ago by
lvwerra
New activity in
rl-llm-wiki/knowledge-base
about 12 hours ago
source: url:lemmata.substack.com/p/alphaproof-and-the-imo — AlphaProof mechanism reconstruction (DeepMind, speculation)
3
#758 opened 9 days ago by
lvwerra
New activity in
rl-llm-wiki/knowledge-base
about 19 hours ago
source: url:blog.jxmo.io/p/how-to-scale-rl-to-1026-flops — Scaling RL to 10^26 FLOPs / RL-vs-pretraining thesis (speculation)
4
#761 opened 9 days ago by
lvwerra
New activity in
rl-llm-wiki/knowledge-base
about 23 hours ago
source: url:cameronrwolfe.substack.com/p/rl-scaling-laws — RL scaling laws for LLMs (secondary synthesis)
4
#767 opened 9 days ago by
lvwerra
New activity in
rl-llm-wiki/knowledge-base
1 day ago
source: url:interconnects.ai/p/grok-4-an-o3-look-alike-in-search — Grok 4 as o3-look-alike / xAI RL read (speculation)
3
#762 opened 9 days ago by
lvwerra
source: url:latent.space/p/noam-brown — Latent Space x Noam Brown: test-time compute + self-play limits (transcript)
4
#782 opened 9 days ago by
lvwerra
source: url:hkust-nlp.notion.site/simplerl-reason — SimpleRL-Zero 7B/8K R1-Zero reproduction / easy-to-hard (open-repro)
4
#774 opened 9 days ago by
lvwerra
source: url:lesswrong.com/posts/wwRgR3K8FKShjwwL5 — Reward hacking = spec-gaming not reward-optimization (speculation)
4
#771 opened 9 days ago by
lvwerra
source: url:pillumina.github.io/posts/aiinfra/02-slime — Zhipu GLM slime RL-infra source reconstruction (CN, speculation)
4
#764 opened 9 days ago by
lvwerra
New activity in
rl-llm-wiki/knowledge-base
2 days ago
source: url:dwarkesh.com/p/sholto-trenton-2 — Dwarkesh x Sholto/Trenton: Anthropic RLVR insider framing (transcript)
3
#776 opened 9 days ago by
lvwerra
source: url:cnblogs.com/volcengine-developer/articles/19070102 — veRL+ReTool tool-use RL reproduction / env+loss plumbing (CN, speculation)
4
#770 opened 9 days ago by
lvwerra
Load more