Sitemap
A list of all the posts and pages found on the site. For you robots out there, there is an XML version available for digesting as well.
Pages
Posts
portfolio
publications
Triad: A Framework Leveraging a Multi-Role LLM-based Agent to Solve Knowledge Base Question Answering
A multi-role LLM agent framework for knowledge base question answering.
OpenReview ›S³cMath: Spontaneous Step-Level Self-Correction Makes Large Language Models Better Mathematical Reasoners
Spontaneous step-level self-correction makes LLMs stronger mathematical reasoners.
OpenReview ›LogicPro: Improving Complex Logical Reasoning via Program-Guided Learning
Improves complex logical reasoning by learning from program-guided supervision.
OpenReview ›UGPhysics: A Comprehensive Benchmark for Undergraduate Physics Reasoning with Large Language Models
A comprehensive benchmark for undergraduate-level physics reasoning in LLMs.
OpenReview ›Cooper: Co-Optimizing Policy and Reward Models in Reinforcement Learning for Large Language Models
Co-optimizes the policy and reward models together in reinforcement learning for LLMs.
arXiv ›GSM8K-V: Can Vision Language Models Solve Grade School Math Word Problems in Visual Contexts
Tests whether vision-language models can solve grade-school math word problems shown visually.
arXiv ›SVGenius: Benchmarking LLMs in SVG Understanding, Editing and Generation
A benchmark for LLM understanding, editing, and generation of SVG graphics.
OpenReview ›Every Step Evolves: Scaling Reinforcement Learning for Trillion-Scale Thinking Model (Ring-1T)
Technical report on scaling reinforcement learning to a trillion-parameter reasoning model.
arXiv ›AskToAct: Enhancing LLMs Tool Use via Self-Correcting Clarification
Enhances LLM tool use by self-correcting clarification of ambiguous requests.
OpenReview ›Do Large Language Models excel in Complex Logical Reasoning with Formal Language?
Investigates whether LLMs can perform complex logical reasoning expressed in formal language.
OpenReview ›Mind the Gap: Bridging Thought Leap for Improved Chain-of-Thought Tuning
Detects and bridges thought leaps in chains of thought to improve reasoning fine-tuning.
OpenReview ›Let LRMs Break Free from Overthinking via Self-Braking Tuning
Teaches reasoning models to stop overthinking by learning when to brake.
OpenReview ›Test-Time Reinforcement Learning for GUI Grounding via Region Consistency
Improves GUI grounding at test time via label-free region-consistency reinforcement learning.
OpenReview ›Teaching LLMs According to Their Aptitude: Adaptive Switching Between CoT and TIR for Mathematical Problem Solving
Adaptively switches between chain-of-thought and tool-integrated reasoning based on problem difficulty.
OpenReview ›OmniEAR: Benchmarking Agent Reasoning in Embodied Tasks
Benchmarks agent reasoning in embodied tasks across diverse environments.
OpenReview ›SpatialLadder: Progressive Training for Spatial Reasoning in Vision-Language Models
Progressive training that builds up spatial reasoning ability in vision-language models.
OpenReview ›VerifyBench: Benchmarking Reference-based Reward Systems for Large Language Models
A benchmark for evaluating reference-based reward systems that verify LLM answers.
OpenReview ›MathFimer: Enhancing Mathematical Reasoning by Expanding Reasoning Steps through Fill-in-the-Middle Task
Expands reasoning chains via a fill-in-the-middle task to strengthen mathematical reasoning.
OpenReview ›InftyThink: Breaking the Length Limits of Long-Context Reasoning in Large Language Models
Breaks context-length limits by turning long reasoning into iterative, bounded-length thinking.
OpenReview ›InftyThink+: Effective and Efficient Infinite-Horizon Reasoning via Reinforcement Learning
Scales infinite-horizon reasoning with reinforcement learning for more effective and efficient long thinking.
OpenReview ›Milestone-Guided Policy Learning for Long-Horizon Language Agents
Guides long-horizon language agents with milestone signals to make policy learning more stable.
OpenReview ›