Saved in:
| Main Authors: | Wu, Qiucheng, Zhao, Handong, Saxon, Michael, Bui, Trung, Wang, William Yang, Zhang, Yang, Chang, Shiyu |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | https://arxiv.org/abs/2407.01863 |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Rethinking the Text-Vision Reasoning Imbalance in MLLMs through the Lens of Training Recipes
by: Yao, Guanyu, et al.
Published: (2025)
by: Yao, Guanyu, et al.
Published: (2025)
VividCam: Learning Unconventional Camera Motions from Virtual Synthetic Videos
by: Wu, Qiucheng, et al.
Published: (2025)
by: Wu, Qiucheng, et al.
Published: (2025)
A Hierarchical Probabilistic Framework for Incremental Knowledge Tracing in Classroom Settings
by: Gao, Xinyi, et al.
Published: (2025)
by: Gao, Xinyi, et al.
Published: (2025)
Losing Visual Needles in Image Haystacks: Vision Language Models are Easily Distracted in Short and Long Contexts
by: Sharma, Aditya, et al.
Published: (2024)
by: Sharma, Aditya, et al.
Published: (2024)
RetouchIQ: MLLM Agents for Instruction-Based Image Retouching with Generalist Reward
by: Wu, Qiucheng, et al.
Published: (2026)
by: Wu, Qiucheng, et al.
Published: (2026)
THOUGHTTERMINATOR: Benchmarking, Calibrating, and Mitigating Overthinking in Reasoning Models
by: Pu, Xiao, et al.
Published: (2025)
by: Pu, Xiao, et al.
Published: (2025)
Valuable Hallucinations: Realizable Non-realistic Propositions
by: Chen, Qiucheng, et al.
Published: (2025)
by: Chen, Qiucheng, et al.
Published: (2025)
Augment before You Try: Knowledge-Enhanced Table Question Answering via Table Expansion
by: Liu, Yujian, et al.
Published: (2024)
by: Liu, Yujian, et al.
Published: (2024)
Benchmarks as Microscopes: A Call for Model Metrology
by: Saxon, Michael, et al.
Published: (2024)
by: Saxon, Michael, et al.
Published: (2024)
Large Language Models Are Latent Variable Models: Explaining and Finding Good Demonstrations for In-Context Learning
by: Wang, Xinyi, et al.
Published: (2023)
by: Wang, Xinyi, et al.
Published: (2023)
Do You Know About My Nation? Investigating Multilingual Language Models' Cultural Literacy Through Factual Knowledge
by: Tanwar, Eshaan, et al.
Published: (2025)
by: Tanwar, Eshaan, et al.
Published: (2025)
Prism: A Framework for Decoupling and Assessing the Capabilities of VLMs
by: Qiao, Yuxuan, et al.
Published: (2024)
by: Qiao, Yuxuan, et al.
Published: (2024)
CoT-Self-Instruct: Building high-quality synthetic prompts for reasoning and non-reasoning tasks
by: Yu, Ping, et al.
Published: (2025)
by: Yu, Ping, et al.
Published: (2025)
Lost in Translation? Translation Errors and Challenges for Fair Assessment of Text-to-Image Models on Multilingual Concepts
by: Saxon, Michael, et al.
Published: (2024)
by: Saxon, Michael, et al.
Published: (2024)
Who Evaluates the Evaluations? Objectively Scoring Text-to-Image Prompt Coherence Metrics with T2IScoreScore (TS2)
by: Saxon, Michael, et al.
Published: (2024)
by: Saxon, Michael, et al.
Published: (2024)
DynaSaur: Large Language Agents Beyond Predefined Actions
by: Nguyen, Dang, et al.
Published: (2024)
by: Nguyen, Dang, et al.
Published: (2024)
TC-Bench: Benchmarking Temporal Compositionality in Text-to-Video and Image-to-Video Generation
by: Feng, Weixi, et al.
Published: (2024)
by: Feng, Weixi, et al.
Published: (2024)
Have the VLMs Lost Confidence? A Study of Sycophancy in VLMs
by: Li, Shuo, et al.
Published: (2024)
by: Li, Shuo, et al.
Published: (2024)
Code-enabled language models can outperform reasoning models on diverse tasks
by: Zhang, Cedegao E., et al.
Published: (2025)
by: Zhang, Cedegao E., et al.
Published: (2025)
Fictitious Synthetic Data Can Improve LLM Factuality via Prerequisite Learning
by: Liu, Yujian, et al.
Published: (2024)
by: Liu, Yujian, et al.
Published: (2024)
Revisiting Who's Harry Potter: Towards Targeted Unlearning from a Causal Intervention Perspective
by: Liu, Yujian, et al.
Published: (2024)
by: Liu, Yujian, et al.
Published: (2024)
A Probabilistic Framework for LLM Hallucination Detection via Belief Tree Propagation
by: Hou, Bairu, et al.
Published: (2024)
by: Hou, Bairu, et al.
Published: (2024)
QuestBench: Can LLMs ask the right question to acquire information in reasoning tasks?
by: Li, Belinda Z., et al.
Published: (2025)
by: Li, Belinda Z., et al.
Published: (2025)
Language models show human-like content effects on reasoning tasks
by: Dasgupta, Ishita, et al.
Published: (2022)
by: Dasgupta, Ishita, et al.
Published: (2022)
Superhuman performance of a large language model on the reasoning tasks of a physician
by: Brodeur, Peter G., et al.
Published: (2024)
by: Brodeur, Peter G., et al.
Published: (2024)
Bridging the Reasoning Gap in Vietnamese with Small Language Models via Test-Time Scaling
by: Trung, Bui The, et al.
Published: (2026)
by: Trung, Bui The, et al.
Published: (2026)
Evidence from counterfactual tasks supports emergent analogical reasoning in large language models
by: Webb, Taylor, et al.
Published: (2024)
by: Webb, Taylor, et al.
Published: (2024)
Discovering Low-rank Subspaces for Language-agnostic Multilingual Representations
by: Xie, Zhihui, et al.
Published: (2024)
by: Xie, Zhihui, et al.
Published: (2024)
AKEW: Assessing Knowledge Editing in the Wild
by: Wu, Xiaobao, et al.
Published: (2024)
by: Wu, Xiaobao, et al.
Published: (2024)
CORG: Generating Answers from Complex, Interrelated Contexts
by: Lee, Hyunji, et al.
Published: (2025)
by: Lee, Hyunji, et al.
Published: (2025)
Reinforced Curriculum Pre-Alignment for Domain-Adaptive VLMs
by: Yan, Yuming, et al.
Published: (2026)
by: Yan, Yuming, et al.
Published: (2026)
The Eloquence team submission for task 1 of MLC-SLM challenge
by: Concina, Lorenzo, et al.
Published: (2025)
by: Concina, Lorenzo, et al.
Published: (2025)
Reminding Multimodal Large Language Models of Object-aware Knowledge with Retrieved Tags
by: Qi, Daiqing, et al.
Published: (2024)
by: Qi, Daiqing, et al.
Published: (2024)
Online-PVLM: Advancing Personalized VLMs with Online Concept Learning
by: Bai, Huiyu, et al.
Published: (2025)
by: Bai, Huiyu, et al.
Published: (2025)
Distributional reasoning in LLMs: Parallel reasoning processes in multi-hop reasoning
by: Shalev, Yuval, et al.
Published: (2024)
by: Shalev, Yuval, et al.
Published: (2024)
CAIRe: Cultural Attribution of Images by Retrieval-Augmented Evaluation
by: Yayavaram, Arnav, et al.
Published: (2025)
by: Yayavaram, Arnav, et al.
Published: (2025)
Culture is Everywhere: A Call for Intentionally Cultural Evaluation
by: Oh, Juhyun, et al.
Published: (2025)
by: Oh, Juhyun, et al.
Published: (2025)
GUI-Bee: Align GUI Action Grounding to Novel Environments via Autonomous Exploration
by: Fan, Yue, et al.
Published: (2025)
by: Fan, Yue, et al.
Published: (2025)
Aviary: training language agents on challenging scientific tasks
by: Narayanan, Siddharth, et al.
Published: (2024)
by: Narayanan, Siddharth, et al.
Published: (2024)
How Well Do Agentic Skills Work in the Wild: Benchmarking LLM Skill Usage in Realistic Settings
by: Liu, Yujian, et al.
Published: (2026)
by: Liu, Yujian, et al.
Published: (2026)
Similar Items
-
Rethinking the Text-Vision Reasoning Imbalance in MLLMs through the Lens of Training Recipes
by: Yao, Guanyu, et al.
Published: (2025) -
VividCam: Learning Unconventional Camera Motions from Virtual Synthetic Videos
by: Wu, Qiucheng, et al.
Published: (2025) -
A Hierarchical Probabilistic Framework for Incremental Knowledge Tracing in Classroom Settings
by: Gao, Xinyi, et al.
Published: (2025) -
Losing Visual Needles in Image Haystacks: Vision Language Models are Easily Distracted in Short and Long Contexts
by: Sharma, Aditya, et al.
Published: (2024) -
RetouchIQ: MLLM Agents for Instruction-Based Image Retouching with Generalist Reward
by: Wu, Qiucheng, et al.
Published: (2026)