Tables as Images? Exploring the Strengths and Limitations of LLMs on Multimodal Representations of Tabular Data Paper • 2402.12424 • Published Feb 19, 2024
The Wrong Kind of Right: Quantifying and Localizing Misfired Alignment in LLMs Paper • 2606.18656 • Published Jun 17
Can LLM Agents Stick to the Script? A Benchmark for Long-Horizon Consistency in Interactive Narratives Paper • 2608.08160 • Published Aug 8 • 30
Recent Advances in Text-to-SQL: A Survey of What We Have and What We Expect Paper • 2208.10099 • Published Aug 22, 2022
HI-TOM: A Benchmark for Evaluating Higher-Order Theory of Mind Reasoning in Large Language Models Paper • 2310.16755 • Published Oct 25, 2023
Confidence Comes from Experience: Experiential Confidence Estimation from Reasoning to Agents Paper • 2609.17708 • Published 23 days ago • 78
Revisiting Cross-Lingual Summarization: A Corpus-based Study and A New Benchmark with Improved Annotation Paper • 2307.04018 • Published Jul 8, 2023
UniSumm and SummZoo: Unified Model and Diverse Benchmark for Few-Shot Summarization Paper • 2211.09783 • Published Nov 17, 2022
See What LLMs Cannot Answer: A Self-Challenge Framework for Uncovering LLM Weaknesses Paper • 2408.08978 • Published Aug 16, 2024
Siren's Song in the AI Ocean: A Survey on Hallucination in Large Language Models Paper • 2309.01219 • Published Sep 3, 2023 • 2
The Gold Medals in an Empty Room: Diagnosing Metalinguistic Reasoning in LLMs with Camlang Paper • 2509.00425 • Published Aug 30, 2025 • 12
Next Token Prediction Towards Multimodal Intelligence: A Comprehensive Survey Paper • 2412.18619 • Published Dec 16, 2024 • 60
DialogSum: A Real-Life Scenario Dialogue Summarization Dataset Paper • 2105.06762 • Published May 14, 2021