Hugging Face
Models
Datasets
Spaces
Buckets
new
Docs
Enterprise
Pricing
Website
Tasks
HuggingChat
Collections
Languages
Organizations
Community
Blog
Posts
Daily Papers
Hardware
Learn
Discord
Forum
GitHub
Solutions
Team & Enterprise
Hugging Face PRO
Enterprise Support
Inference Providers
Inference Endpoints
Storage Buckets
Log In
Sign Up
🔄
In a Training Loop
Quentin Gallouédec
PRO
qgallouedec
1762
186
147
Follow
wubro's profile picture
thcorre's profile picture
amir1371's profile picture
694 followers
·
370 following
QGallouedec
qgallouedec
qgallouedec
qgallouedec.bsky.social
AI & ML interests
None yet
Recent Activity
updated
a bucket
about 19 hours ago
qgallouedec/cron
updated
a bucket
1 day ago
hf-doc-build/doc-dev
reacted
to
sergiopaniego
's
post
with 👍
2 days ago
most multi-turn RL loops have a silent bug: you decode the model's output to detect tool calls, then re-tokenize the conversation for the next turn. BPE isn't invertible, so decode then re-encode can land on different ids. gradient ends up on tokens the model never sampled. no crash, just quietly wrong math and broken training @qgallouedec wrote a super educational blog on MITO (message-in, token-out) vs TITO (token-in, token-out) and how you might fix the problem above go read it 🤓 https://qgallouedec-tito.hf.space/
View all activity
Organizations
qgallouedec
's datasets
88
Sort: Recently updated
qgallouedec/tool-calls-mini
Viewer
•
Updated
30 days ago
•
500
•
79
qgallouedec/one-line-answers
Viewer
•
Updated
about 1 month ago
•
8.19k
•
83
qgallouedec/guess-the-regex
Viewer
•
Updated
Jun 21
•
213
•
3.85k
qgallouedec/test-grpo-vlm-log-completions
Viewer
•
Updated
Mar 20
•
435
•
672
qgallouedec/llama_star_formatted
Viewer
•
Updated
Feb 21
•
7.21k
•
22
qgallouedec/deepmath-completions-logs2
Viewer
•
Updated
Jan 22
•
48
•
94
qgallouedec/deepmath-completions-logs
Viewer
•
Updated
Jan 13
•
232
•
945
•
1
qgallouedec/Dolci-Think-DPO-7B
Viewer
•
Updated
Nov 28, 2025
•
150k
•
32
qgallouedec/biogrid_qa
Viewer
•
Updated
Nov 18, 2025
•
59.4k
•
202
qgallouedec/human_gene_interaction_qa_v2
Viewer
•
Updated
Nov 18, 2025
•
79.2k
•
31
qgallouedec/human_gene_interaction_qa
Viewer
•
Updated
Nov 17, 2025
•
1.84M
•
38
qgallouedec/biogrid
Viewer
•
Updated
Nov 17, 2025
•
2.82M
•
382
qgallouedec/trl-metrics
Viewer
•
Updated
Oct 7, 2025
•
148k
•
234
•
1
qgallouedec/rick
Viewer
•
Updated
Sep 11, 2025
•
1.18k
•
23
qgallouedec/OpenMathReasoning
Viewer
•
Updated
Sep 10, 2025
•
10k
•
16
qgallouedec/math-lvl3to5-8k
Viewer
•
Updated
Aug 22, 2025
•
8.52k
•
24
qgallouedec/svg
Viewer
•
Updated
Aug 2, 2025
•
900
•
29
•
1
qgallouedec/rick-physics-grpo
Viewer
•
Updated
May 22, 2025
•
1.79k
•
29
•
1
qgallouedec/rick-science
Viewer
•
Updated
May 16, 2025
•
1.18k
•
19
•
3
qgallouedec/physics-problems
Viewer
•
Updated
May 10, 2025
•
247
•
69
•
1
qgallouedec/rick-teaches-math
Viewer
•
Updated
May 10, 2025
•
6.8k
•
22
qgallouedec/DAPO-Math-17k-Processed-Scored
Viewer
•
Updated
Apr 29, 2025
•
16.4k
•
37
•
3
qgallouedec/prm800k
Viewer
•
Updated
Dec 17, 2024
•
41.2k
•
27
•
3
qgallouedec/ultrafeedback-prompt
Viewer
•
Updated
Sep 9, 2024
•
60.9k
•
30
qgallouedec/ultrafeedback-gpt-3.5-turbo-helpfulness
Viewer
•
Updated
Sep 9, 2024
•
16.6k
•
15
qgallouedec/lm-human-preferences-descriptiveness
Viewer
•
Updated
Sep 9, 2024
•
6.26k
•
19
qgallouedec/lm-human-preferences-sentiment
Viewer
•
Updated
Sep 9, 2024
•
6.26k
•
21
qgallouedec/tldr-preference
Viewer
•
Updated
Sep 9, 2024
•
179k
•
25
qgallouedec/tldr
Viewer
•
Updated
Sep 9, 2024
•
130k
•
31
qgallouedec/hh-rlhf-helpful-base
Viewer
•
Updated
Sep 5, 2024
•
46.2k
•
24
Previous
1
2
3
Next