WebArXiv: Evaluating Multimodal Agents on Time-Invariant arXiv Tasks
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Sun, Zihao, Chen, Ling |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Documentation Retrieval Improves Planning Language Generation
von: Wang, Renxiang, et al.
Veröffentlicht: (2025)
von: Wang, Renxiang, et al.
Veröffentlicht: (2025)
Leveraging Translation For Optimal Recall: Tailoring LLM Personalization With User Profiles
von: Ravichandran, Karthik, et al.
Veröffentlicht: (2024)
von: Ravichandran, Karthik, et al.
Veröffentlicht: (2024)
Evaluating the efficacy of LLM Safety Solutions : The Palit Benchmark Dataset
von: Palit, Sayon, et al.
Veröffentlicht: (2025)
von: Palit, Sayon, et al.
Veröffentlicht: (2025)
PromptSAM+: Malware Detection based on Prompt Segment Anything Model
von: Wei, Xingyuan, et al.
Veröffentlicht: (2024)
von: Wei, Xingyuan, et al.
Veröffentlicht: (2024)
Circularity and Symmetries of $p$ and $p^{2}$-polygons
von: Haag, Rolf
Veröffentlicht: (2025)
von: Haag, Rolf
Veröffentlicht: (2025)
All for law and law for all: Adaptive RAG Pipeline for Legal Research
von: Keisha, Figarri, et al.
Veröffentlicht: (2025)
von: Keisha, Figarri, et al.
Veröffentlicht: (2025)
Mobile Phone Sensor-based Nigerian Driving Dataset to Detect Alcohol-influenced Behaviours
von: Thompson, Iniakpokeikiye Peter, et al.
Veröffentlicht: (2025)
von: Thompson, Iniakpokeikiye Peter, et al.
Veröffentlicht: (2025)
Random Heterogeneous Neurochaos Learning Architecture for Data Classification
von: S, Remya Ajai A, et al.
Veröffentlicht: (2024)
von: S, Remya Ajai A, et al.
Veröffentlicht: (2024)
The Hidden Attention of Mamba Models
von: Ali, Ameen, et al.
Veröffentlicht: (2024)
von: Ali, Ameen, et al.
Veröffentlicht: (2024)
Personalized Federated Sequential Recommender
von: Di, Yicheng
Veröffentlicht: (2026)
von: Di, Yicheng
Veröffentlicht: (2026)
From Pixels to Privacy: Temporally Consistent Video Anonymization via Token Pruning for Privacy Preserving Action Recognition
von: Aslam, Nazia, et al.
Veröffentlicht: (2026)
von: Aslam, Nazia, et al.
Veröffentlicht: (2026)
Automotive-ENV: Benchmarking Multimodal Agents in Vehicle Interface Systems
von: Yan, Junfeng, et al.
Veröffentlicht: (2025)
von: Yan, Junfeng, et al.
Veröffentlicht: (2025)
The Evolution of Embedding Table Optimization and Multi-Epoch Training in Pinterest Ads Conversion
von: Qiu, Andrew, et al.
Veröffentlicht: (2025)
von: Qiu, Andrew, et al.
Veröffentlicht: (2025)
I-WebGenBench : Evaluating Interactivity in LLM-Generated Scientific Web Applications
von: Dai, Dasen, et al.
Veröffentlicht: (2026)
von: Dai, Dasen, et al.
Veröffentlicht: (2026)
DataFactory: Collaborative Multi-Agent Framework for Advanced Table Question Answering
von: Wang, Tong, et al.
Veröffentlicht: (2026)
von: Wang, Tong, et al.
Veröffentlicht: (2026)
Software Implementation of Digital Filtering via Tustin's Bilinear Transform
von: Herron, Connor W.
Veröffentlicht: (2024)
von: Herron, Connor W.
Veröffentlicht: (2024)
Detection of ChatGPT Fake Science with the xFakeSci Learning Algorithm
von: Hamed, Ahmed Abdeen, et al.
Veröffentlicht: (2023)
von: Hamed, Ahmed Abdeen, et al.
Veröffentlicht: (2023)
Understanding and Improving Information Preservation in Prompt Compression for LLMs
von: Łajewska, Weronika, et al.
Veröffentlicht: (2025)
von: Łajewska, Weronika, et al.
Veröffentlicht: (2025)
Both Ends Count! Just How Good are LLM Agents at "Text-to-Big SQL"?
von: Eizaguirre, Germán T., et al.
Veröffentlicht: (2026)
von: Eizaguirre, Germán T., et al.
Veröffentlicht: (2026)
R-Genie: Reasoning-Guided Generative Image Editing
von: Zhang, Dong, et al.
Veröffentlicht: (2025)
von: Zhang, Dong, et al.
Veröffentlicht: (2025)
Train to Defend: First Defense Against Cryptanalytic Neural Network Parameter Extraction Attacks
von: Kurian, Ashley, et al.
Veröffentlicht: (2025)
von: Kurian, Ashley, et al.
Veröffentlicht: (2025)
Unraveling the Italian and English Telegram Conspiracy Spheres through Message Forwarding
von: Alvisi, Lorenzo, et al.
Veröffentlicht: (2024)
von: Alvisi, Lorenzo, et al.
Veröffentlicht: (2024)
Towards Platonic Representation for Table Reasoning: A Foundation for Permutation-Invariant Retrieval
von: Tchuitcheu, Willy Carlos, et al.
Veröffentlicht: (2026)
von: Tchuitcheu, Willy Carlos, et al.
Veröffentlicht: (2026)
CBR -- Boosting Adaptive Classification By Retrieval of Encrypted Network Traffic with Out-of-distribution
von: Lukach, Amir, et al.
Veröffentlicht: (2024)
von: Lukach, Amir, et al.
Veröffentlicht: (2024)
PaperVoyager : Building Interactive Web with Visual Language Models
von: Dai, Dasen, et al.
Veröffentlicht: (2026)
von: Dai, Dasen, et al.
Veröffentlicht: (2026)
Only Whats Necessary: Pareto Optimal Data Minimization for Privacy Preserving Video Anomaly Detection
von: Aslam, Nazia, et al.
Veröffentlicht: (2026)
von: Aslam, Nazia, et al.
Veröffentlicht: (2026)
Vision-Language Reasoning for Geolocalization: A Reinforcement Learning Approach
von: Wu, Biao, et al.
Veröffentlicht: (2026)
von: Wu, Biao, et al.
Veröffentlicht: (2026)
A Gamified Evaluation and Recruitment Platform for Low Resource Language Machine Translation Systems
von: Catalan, Carlos Rafael
Veröffentlicht: (2025)
von: Catalan, Carlos Rafael
Veröffentlicht: (2025)
Can Zero-Shot Commercial APIs Deliver Regulatory-Grade Clinical Text DeIdentification?
von: Kocaman, Veysel, et al.
Veröffentlicht: (2025)
von: Kocaman, Veysel, et al.
Veröffentlicht: (2025)
Semantic Caching for Improving Web Affordability
von: Akbar, Hafsa, et al.
Veröffentlicht: (2025)
von: Akbar, Hafsa, et al.
Veröffentlicht: (2025)
Remote Sensing-Oriented World Model
von: Lu, Yuxi, et al.
Veröffentlicht: (2025)
von: Lu, Yuxi, et al.
Veröffentlicht: (2025)
Policy-Grounded Safety Evaluation of 20 Large Language Models
von: Contreras, Juan Manuel
Veröffentlicht: (2025)
von: Contreras, Juan Manuel
Veröffentlicht: (2025)
Infinity Parser: Layout Aware Reinforcement Learning for Scanned Document Parsing
von: Wang, Baode, et al.
Veröffentlicht: (2025)
von: Wang, Baode, et al.
Veröffentlicht: (2025)
Toward Architecture-Aware Evaluation Metrics for LLM Agents
von: Souza, Débora, et al.
Veröffentlicht: (2026)
von: Souza, Débora, et al.
Veröffentlicht: (2026)
Uplink Transmit Power Optimization for Distributed Massive MIMO Systems with 1-Bit ADCs
von: Gouda, Bikshapathi, et al.
Veröffentlicht: (2024)
von: Gouda, Bikshapathi, et al.
Veröffentlicht: (2024)
Segmenting Action-Value Functions Over Time-Scales in SARSA via TD($Δ$)
von: Humayoo, Mahammad
Veröffentlicht: (2024)
von: Humayoo, Mahammad
Veröffentlicht: (2024)
Dreaming Falcon: Physics-Informed Model-Based Reinforcement Learning for Quadcopters
von: Vytla, Eashan, et al.
Veröffentlicht: (2025)
von: Vytla, Eashan, et al.
Veröffentlicht: (2025)
AI-assisted 3D Preservation and Reconstruction of Temple Arts
von: Shih, Naai-Jung
Veröffentlicht: (2025)
von: Shih, Naai-Jung
Veröffentlicht: (2025)
Time-Scale Separation in Q-Learning: Extending TD($\triangle$) for Action-Value Function Decomposition
von: Humayoo, Mahammad
Veröffentlicht: (2024)
von: Humayoo, Mahammad
Veröffentlicht: (2024)
SRSUPM: Sequential Recommender System Based on User Psychological Motivation
von: Di, Yicheng, et al.
Veröffentlicht: (2026)
von: Di, Yicheng, et al.
Veröffentlicht: (2026)
Ähnliche Einträge
-
Documentation Retrieval Improves Planning Language Generation
von: Wang, Renxiang, et al.
Veröffentlicht: (2025) -
Leveraging Translation For Optimal Recall: Tailoring LLM Personalization With User Profiles
von: Ravichandran, Karthik, et al.
Veröffentlicht: (2024) -
Evaluating the efficacy of LLM Safety Solutions : The Palit Benchmark Dataset
von: Palit, Sayon, et al.
Veröffentlicht: (2025) -
PromptSAM+: Malware Detection based on Prompt Segment Anything Model
von: Wei, Xingyuan, et al.
Veröffentlicht: (2024) -
Circularity and Symmetries of $p$ and $p^{2}$-polygons
von: Haag, Rolf
Veröffentlicht: (2025)