Triad: A Framework Leveraging a Multi-Role LLM-based Agent to Solve Knowledge Base Question Answering
A multi-role LLM agent framework for knowledge base question answering.
OpenReview ›A multi-role LLM agent framework for knowledge base question answering.
OpenReview ›Spontaneous step-level self-correction makes LLMs stronger mathematical reasoners.
OpenReview ›Improves complex logical reasoning by learning from program-guided supervision.
OpenReview ›A comprehensive benchmark for undergraduate-level physics reasoning in LLMs.
OpenReview ›Co-optimizes the policy and reward models together in reinforcement learning for LLMs.
arXiv ›Tests whether vision-language models can solve grade-school math word problems shown visually.
arXiv ›A benchmark for LLM understanding, editing, and generation of SVG graphics.
OpenReview ›Technical report on scaling reinforcement learning to a trillion-parameter reasoning model.
arXiv ›Enhances LLM tool use by self-correcting clarification of ambiguous requests.
OpenReview ›Investigates whether LLMs can perform complex logical reasoning expressed in formal language.
OpenReview ›Detects and bridges thought leaps in chains of thought to improve reasoning fine-tuning.
OpenReview ›Teaches reasoning models to stop overthinking by learning when to brake.
OpenReview ›Improves GUI grounding at test time via label-free region-consistency reinforcement learning.
OpenReview ›Adaptively switches between chain-of-thought and tool-integrated reasoning based on problem difficulty.
OpenReview ›Benchmarks agent reasoning in embodied tasks across diverse environments.
OpenReview ›Progressive training that builds up spatial reasoning ability in vision-language models.
OpenReview ›A benchmark for evaluating reference-based reward systems that verify LLM answers.
OpenReview ›Expands reasoning chains via a fill-in-the-middle task to strengthen mathematical reasoning.
OpenReview ›Breaks context-length limits by turning long reasoning into iterative, bounded-length thinking.
OpenReview ›Scales infinite-horizon reasoning with reinforcement learning for more effective and efficient long thinking.
OpenReview ›Guides long-horizon language agents with milestone signals to make policy learning more stable.
OpenReview ›