Tiny Model, Big Logic: Diversity-Driven Optimization Elicits Large-Model Reasoning Ability in VibeThinker-1.5B
-
Updated
Aug 14, 2026 - Python
Tiny Model, Big Logic: Diversity-Driven Optimization Elicits Large-Model Reasoning Ability in VibeThinker-1.5B
[NeurIPS 2025] AGI-Elo: How Far Are We From Mastering A Task?
Benchmark methodology, task sets, and evaluation results for RA²R
Interactive benchmark for evaluating LLMs on clarifying ambiguous code requirements. Paper: arXiv:2607.00711
Resample or reroute after a weak-verifier stop? Pre-registered measurements of recoverable stopping debt on MBPP+, a two-sided action-support gate on BigCodeBench that fails closed, a LiveCodeBench observability ladder, and the exchangeable-actions reference showing realized-maximum gaps carry no selector signal. Artifacts for arXiv:2607.08665v3.
A reproducible LiveCodeBench evaluation scaffold for controlled studies of code-reasoning SFT on Qwen2.5-1.5B.
Self-Improving Agent for Code Generation
LoRA fine-tuning evaluation pipeline for code generation models on LiveCodeBench benchmark
“U.S. Provisional Patent Application No. 63/978,753, filed February 9, 2026” Aleph Generalizable Intelligence Convergence Engine (AGICE) - Evidence-Native, Policy-Governed Reasoning with Rollback and Geometrical Multi-Dimensional Transformation (GMDT) via the Complex Operator (−i)
AI 大模型第三方权威排行榜聚合:LMArena、Artificial Analysis、LiveCodeBench、SuperCLUE、HF Open LLM。无后端纯静态站,GitHub Actions 每日自动抓取更新,含升降箭头/性价比象限/历史趋势。 | Independent AI model leaderboard aggregator - zero-backend static site, auto-updated daily.
NeoSmith Maestro — frontier-level coding accuracy from a bouquet of small models at ~1/30th the cost. Technical report + reproducible evidence (LiveCodeBench v6 92.2%, SWE-bench Pro 74.2/88.9).
Self-hosted LLM evaluation lab: MMLU-Pro / BFCL / IFEval / LiveCodeBench / HumanEval+ runs on Tesla V100 and RTX 3090, llama.cpp serving (incl. Volta sm_70), patched API benchmark clients
Executable-code evaluation and verifier-guided inference, with reproducible benchmarks and separate model-weight retention audits.
To associate your repository with the livecodebench topic, visit your repo's landing page and select "manage topics."