🏗️ Building on HF
Carlos Miguel Patiño
AI & ML interests
None yet
Recent Activity
updated a bucket 1 day ago
sair-distillation/eq2-main-bucket updated a model 1 day ago
cmpatino/qwen-distill-r2 updated a Space 1 day ago
cmpatino/qwen-distill-static-bb06bdOrganizations
topic: dpo-variants runnable SimPO ORPO checks
2
#466 opened 25 days ago
by
cmpatino
source: arxiv:2501.01821 - SDPO
4
#296 opened about 1 month ago
by
cmpatino
topic: algorithms/dpo-variants - add SDPO
2
#297 opened about 1 month ago
by
cmpatino
topic: rl-training-stability-in-practice — weave in PPO-max (Secrets-I) + entropy mechanism
8
#292 opened about 1 month ago
by
hf-dwarez
topic: rollout-generation-infra — colocated resharding engine + generator layout (clean reopen of #271)
5
#290 opened about 1 month ago
by
hf-dwarez
topic: NEW algorithms/self-improvement-and-self-play — method-family hub (STaR/SPIN/Self-Rewarding/Absolute-Zero/TTRL)
2
#286 opened about 1 month ago
by
lvwerra
fix: deepen scalable-oversight §4 with empirical debate, easy→hard, prover-verifier (absorbs 3 orphan sources)
4
#288 opened about 1 month ago
by
lvwerra
topic: iterate rl-for-math-and-code — bring current with the 2025 RLVR wave (recipes, data frontiers, elicit-vs-expand)
4
#277 opened about 1 month ago
by
lvwerra
source: arxiv:2406.11939 — From Crowdsourced Data to High-Quality Benchmarks: Arena-Hard and BenchBuilder Pipeline
4
#259 opened about 1 month ago
by
lvwerra
source: arxiv:2404.01833 — Great, Now Write an Article About That: The Crescendo Multi-Turn LLM Jailbreak Attack
4
#254 opened about 1 month ago
by
lvwerra
topic: iterate preference-reward-models — multi-attribute regression reward models (SteerLM + HelpSteer/HelpSteer2)
4
#252 opened about 1 month ago
by
lvwerra
source: arxiv:2209.00626 — The Alignment Problem from a Deep Learning Perspective
6
#214 opened about 1 month ago
by
lvwerra
source: arxiv:2405.01481 — NeMo-Aligner (clean reopen of #272)
5
#291 opened about 1 month ago
by
hf-dwarez
topic: iterate human-preference-collection — active preference learning / query efficiency (APRIL)
2
#284 opened about 1 month ago
by
lvwerra
topic: iterate ai-feedback-data — UltraFeedback dataset, RLAIF head-to-head, RLAIF-V open-MLLM feedback
2
#283 opened about 1 month ago
by
lvwerra
topic: iterate verifiable-rewards — attribution caveat: how load-bearing is the verifier's correctness?
2
#282 opened about 1 month ago
by
lvwerra
topic: iterate data-quality-and-filtering — Skywork-Reward (quality>scale, decontam) + HelpSteer2 annotation QA
2
#281 opened about 1 month ago
by
lvwerra
topic: iterate rlvr-overview — complete §5 with the 2025 elicit-vs-expand evidence
2
#280 opened about 1 month ago
by
lvwerra
topic: iterate rlaif — RLAIF-V (open AI feedback + self-alignment for multimodal models)
2
#279 opened about 1 month ago
by
lvwerra
topic: iterate reward-hacking — reward tampering + frontier verifier hacking + CoT-monitoring (and its fragility)
2
#278 opened about 1 month ago
by
lvwerra