Space for RePro: Training Language Models to Faithfully Recycle the Web for Pretraining
AI & ML interests
None defined yet.
Recent Activity
View all activity
Papers
EmbodiedMidtrain: Bridging the Gap between Vision-Language Models and Vision-Language-Action Models via Mid-training
RePro: Training Language Models to Faithfully Recycle the Web for Pretraining
models 5
cx-cmu/AutoGEO_mini_Qwen1.7B_ResearchyGEO
Text Generation • 2B • Updated • 24 •
cx-cmu/AutoGEO_mini_Qwen1.7B_GEOBench
Text Generation • 2B • Updated • 20 •
cx-cmu/AutoGEO_mini_Qwen1.7B_Ecommerce
Text Generation • 2B • Updated • 24 •
cx-cmu/repro-rephraser-4B
Text Generation • 196k • Updated • 54 • • 2
cx-cmu/repro-rephraser-1B
Text Generation • 1B • Updated • 5
datasets 10
cx-cmu/agent_trajectories
Updated • 209 • 1
cx-cmu/deepresearchgym-agentic-search-logs
Viewer • Updated • 14.3M • 315 • 14
cx-cmu/Researchy-GEO
Viewer • Updated • 47k • 846 • 1
cx-cmu/GEO-Bench
Viewer • Updated • 37.4k • 133 • 1
cx-cmu/E-commerce
Viewer • Updated • 7.97k • 472 • 2
cx-cmu/ClueWeb-Reco
Viewer • Updated • 87.2M • 70 • 1
cx-cmu/repro-organic-data-72B
Viewer • Updated • 58.3M • 1.04k
cx-cmu/repro-rl-data
Viewer • Updated • 41k • 71
cx-cmu/repro-rephrased-data-72B
Viewer • Updated • 39M • 446
cx-cmu/CLUE-LLM
Viewer • Updated • 1.21k • 10