Gespeichert in:
| Hauptverfasser: | Abaskohi, Amirhossein, He, Yuhang, West, Peter, Carenini, Giuseppe, Chawla, Pranit, Vineet, Vibhav |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | https://arxiv.org/abs/2605.11212 |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
BCAmirs at SemEval-2024 Task 4: Beyond Words: A Multimodal and Multilingual Exploration of Persuasion in Memes
von: Abaskohi, Amirhossein, et al.
Veröffentlicht: (2024)
von: Abaskohi, Amirhossein, et al.
Veröffentlicht: (2024)
CEMTM: Contextual Embedding-based Multimodal Topic Modeling
von: Abaskohi, Amirhossein, et al.
Veröffentlicht: (2025)
von: Abaskohi, Amirhossein, et al.
Veröffentlicht: (2025)
Improving Neural Topic Modeling with Semantically-Grounded Soft Label Distributions
von: Li, Raymond, et al.
Veröffentlicht: (2026)
von: Li, Raymond, et al.
Veröffentlicht: (2026)
FM2DS: Few-Shot Multimodal Multihop Data Synthesis with Knowledge Distillation for Question Answering
von: Abaskohi, Amirhossein, et al.
Veröffentlicht: (2024)
von: Abaskohi, Amirhossein, et al.
Veröffentlicht: (2024)
AgentAda: Skill-Adaptive Data Analytics for Tailored Insight Discovery
von: Abaskohi, Amirhossein, et al.
Veröffentlicht: (2025)
von: Abaskohi, Amirhossein, et al.
Veröffentlicht: (2025)
ChartGaze: Enhancing Chart Understanding in LVLMs with Eye-Tracking Guided Attention Refinement
von: Salamatian, Ali, et al.
Veröffentlicht: (2025)
von: Salamatian, Ali, et al.
Veröffentlicht: (2025)
Infusing Theory of Mind into Socially Intelligent LLM Agents
von: Hwang, EunJeong, et al.
Veröffentlicht: (2025)
von: Hwang, EunJeong, et al.
Veröffentlicht: (2025)
PyramidDrop: Accelerating Your Large Vision-Language Models via Pyramid Visual Redundancy Reduction
von: Xing, Long, et al.
Veröffentlicht: (2024)
von: Xing, Long, et al.
Veröffentlicht: (2024)
ReVision: A Dataset and Baseline VLM for Privacy-Preserving Task-Oriented Visual Instruction Rewriting
von: Mishra, Abhijit, et al.
Veröffentlicht: (2025)
von: Mishra, Abhijit, et al.
Veröffentlicht: (2025)
WebSTAR: Scalable Data Synthesis for Computer Use Agents with Step-Level Filtering
von: He, Yifei, et al.
Veröffentlicht: (2025)
von: He, Yifei, et al.
Veröffentlicht: (2025)
uTeBC-NLP at SemEval-2024 Task 9: Can LLMs be Lateral Thinkers?
von: Sadeghi, Pouya, et al.
Veröffentlicht: (2024)
von: Sadeghi, Pouya, et al.
Veröffentlicht: (2024)
OS-Marathon: Benchmarking Computer-Use Agents on Long-Horizon Repetitive Tasks
von: Wu, Jing, et al.
Veröffentlicht: (2026)
von: Wu, Jing, et al.
Veröffentlicht: (2026)
Just Do It!? Computer-Use Agents Exhibit Blind Goal-Directedness
von: Shayegani, Erfan, et al.
Veröffentlicht: (2025)
von: Shayegani, Erfan, et al.
Veröffentlicht: (2025)
DRBench: A Realistic Benchmark for Enterprise Deep Research
von: Abaskohi, Amirhossein, et al.
Veröffentlicht: (2025)
von: Abaskohi, Amirhossein, et al.
Veröffentlicht: (2025)
Exposing the Achilles' Heel: Evaluating LLMs Ability to Handle Mistakes in Mathematical Reasoning
von: Singh, Joykirat, et al.
Veröffentlicht: (2024)
von: Singh, Joykirat, et al.
Veröffentlicht: (2024)
ReVisioning the Public Library as an Oasis of Learning
von: Cassell, Mary A., et al.
Veröffentlicht: (2012)
von: Cassell, Mary A., et al.
Veröffentlicht: (2012)
Captioning Visualizations with Large Language Models (CVLLM): A Tutorial
von: Carenini, Giuseppe, et al.
Veröffentlicht: (2024)
von: Carenini, Giuseppe, et al.
Veröffentlicht: (2024)
A Large-Scale Analysis of Persian Tweets Regarding Covid-19 Vaccination
von: ShabaniMirzaei, Taha, et al.
Veröffentlicht: (2023)
von: ShabaniMirzaei, Taha, et al.
Veröffentlicht: (2023)
BeDiscovER: The Benchmark of Discourse Understanding in the Era of Reasoning Language Models
von: Li, Chuyuan, et al.
Veröffentlicht: (2025)
von: Li, Chuyuan, et al.
Veröffentlicht: (2025)
Personalized Abstractive Summarization by Tri-agent Generation Pipeline
von: Xiao, Wen, et al.
Veröffentlicht: (2023)
von: Xiao, Wen, et al.
Veröffentlicht: (2023)
What MLLMs Learn about When they Learn about Multimodal Reasoning
von: Chung, Jiwan, et al.
Veröffentlicht: (2025)
von: Chung, Jiwan, et al.
Veröffentlicht: (2025)
Improving Language Models with Intentional Analysis
von: Yin, Yuwei, et al.
Veröffentlicht: (2025)
von: Yin, Yuwei, et al.
Veröffentlicht: (2025)
ReVision: Refining Video Diffusion with Explicit 3D Motion Modeling
von: Liu, Qihao, et al.
Veröffentlicht: (2025)
von: Liu, Qihao, et al.
Veröffentlicht: (2025)
Fara-7B: An Efficient Agentic Model for Computer Use
von: Awadallah, Ahmed, et al.
Veröffentlicht: (2025)
von: Awadallah, Ahmed, et al.
Veröffentlicht: (2025)
Neural Multimodal Topic Modeling: A Comprehensive Evaluation
von: González-Pizarro, Felipe, et al.
Veröffentlicht: (2024)
von: González-Pizarro, Felipe, et al.
Veröffentlicht: (2024)
Adaptive Vision-Language Model Routing for Computer Use Agents
von: Liu, Xunzhuo, et al.
Veröffentlicht: (2026)
von: Liu, Xunzhuo, et al.
Veröffentlicht: (2026)
Delta-KNN: Improving Demonstration Selection in In-Context Learning for Alzheimer's Disease Detection
von: Li, Chuyuan, et al.
Veröffentlicht: (2025)
von: Li, Chuyuan, et al.
Veröffentlicht: (2025)
Topic-Guided Reinforcement Learning with LLMs for Enhancing Multi-Document Summarization
von: Li, Chuyuan, et al.
Veröffentlicht: (2025)
von: Li, Chuyuan, et al.
Veröffentlicht: (2025)
Scaling Agents for Computer Use
von: Gonzalez-Pumariega, Gonzalo, et al.
Veröffentlicht: (2025)
von: Gonzalez-Pumariega, Gonzalo, et al.
Veröffentlicht: (2025)
IntentGrasp: A Comprehensive Benchmark for Intent Understanding
von: Yin, Yuwei, et al.
Veröffentlicht: (2026)
von: Yin, Yuwei, et al.
Veröffentlicht: (2026)
ReTaKe: Reducing Temporal and Knowledge Redundancy for Long Video Understanding
von: Wang, Xiao, et al.
Veröffentlicht: (2024)
von: Wang, Xiao, et al.
Veröffentlicht: (2024)
RiTTA: Modeling Event Relations in Text-to-Audio Generation
von: He, Yuhang, et al.
Veröffentlicht: (2024)
von: He, Yuhang, et al.
Veröffentlicht: (2024)
Multi2: Multi-Agent Test-Time Scalable Framework for Multi-Document Processing
von: Cao, Juntai, et al.
Veröffentlicht: (2025)
von: Cao, Juntai, et al.
Veröffentlicht: (2025)
$R^2$-dLLM: Accelerating Diffusion Large Language Models via Spatio-Temporal Redundancy Reduction
von: Du, Zhenbang, et al.
Veröffentlicht: (2026)
von: Du, Zhenbang, et al.
Veröffentlicht: (2026)
Verbosity-Aware Rationale Reduction: Effective Reduction of Redundant Rationale via Principled Criteria
von: Jang, Joonwon, et al.
Veröffentlicht: (2024)
von: Jang, Joonwon, et al.
Veröffentlicht: (2024)
SWI: Speaking with Intent in Large Language Models
von: Yin, Yuwei, et al.
Veröffentlicht: (2025)
von: Yin, Yuwei, et al.
Veröffentlicht: (2025)
Physics Knowledge in Frontier Models: A Diagnostic Study of Failure Modes
von: Bagdonaviciute, Ieva, et al.
Veröffentlicht: (2025)
von: Bagdonaviciute, Ieva, et al.
Veröffentlicht: (2025)
AdaReTaKe: Adaptive Redundancy Reduction to Perceive Longer for Video-language Understanding
von: Wang, Xiao, et al.
Veröffentlicht: (2025)
von: Wang, Xiao, et al.
Veröffentlicht: (2025)
Efficient Agent Training for Computer Use
von: He, Yanheng, et al.
Veröffentlicht: (2025)
von: He, Yanheng, et al.
Veröffentlicht: (2025)
A11y-Compressor: A Framework for Enhancing the Efficiency of GUI Agent Observations through Visual Context Reconstruction and Redundancy Reduction
von: Takeshita, Michito, et al.
Veröffentlicht: (2026)
von: Takeshita, Michito, et al.
Veröffentlicht: (2026)
Ähnliche Einträge
-
BCAmirs at SemEval-2024 Task 4: Beyond Words: A Multimodal and Multilingual Exploration of Persuasion in Memes
von: Abaskohi, Amirhossein, et al.
Veröffentlicht: (2024) -
CEMTM: Contextual Embedding-based Multimodal Topic Modeling
von: Abaskohi, Amirhossein, et al.
Veröffentlicht: (2025) -
Improving Neural Topic Modeling with Semantically-Grounded Soft Label Distributions
von: Li, Raymond, et al.
Veröffentlicht: (2026) -
FM2DS: Few-Shot Multimodal Multihop Data Synthesis with Knowledge Distillation for Question Answering
von: Abaskohi, Amirhossein, et al.
Veröffentlicht: (2024) -
AgentAda: Skill-Adaptive Data Analytics for Tailored Insight Discovery
von: Abaskohi, Amirhossein, et al.
Veröffentlicht: (2025)