Building LLM Agents by Incorporating Insights from Computer Systems
Fuente:
arXiv
Salvato in:
| Autori principali: | Mi, Yapeng, Gao, Zhi, Ma, Xiaojian, Li, Qing |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2025
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Iterative Tool Usage Exploration for Multimodal Agents via Step-wise Preference Tuning
di: Li, Pengxiang, et al.
Pubblicazione: (2025)
di: Li, Pengxiang, et al.
Pubblicazione: (2025)
VideoAgent: A Memory-augmented Multimodal Agent for Video Understanding
di: Fan, Yue, et al.
Pubblicazione: (2024)
di: Fan, Yue, et al.
Pubblicazione: (2024)
Multi-modal Agent Tuning: Building a VLM-Driven Agent for Efficient Tool Usage
di: Gao, Zhi, et al.
Pubblicazione: (2024)
di: Gao, Zhi, et al.
Pubblicazione: (2024)
TongUI: Internet-Scale Trajectories from Multimodal Web Tutorials for Generalized GUI Agents
di: Zhang, Bofei, et al.
Pubblicazione: (2025)
di: Zhang, Bofei, et al.
Pubblicazione: (2025)
MILR: Improving Multimodal Image Generation via Test-Time Latent Reasoning
di: Mi, Yapeng, et al.
Pubblicazione: (2025)
di: Mi, Yapeng, et al.
Pubblicazione: (2025)
CLOVA: A Closed-Loop Visual Assistant with Tool Usage and Update
di: Gao, Zhi, et al.
Pubblicazione: (2023)
di: Gao, Zhi, et al.
Pubblicazione: (2023)
Embodied VideoAgent: Persistent Memory from Egocentric Videos and Embodied Sensors Enables Dynamic Scene Understanding
di: Fan, Yue, et al.
Pubblicazione: (2024)
di: Fan, Yue, et al.
Pubblicazione: (2024)
NEP: Autoregressive Image Editing via Next Editing Token Prediction
di: Wu, Huimin, et al.
Pubblicazione: (2025)
di: Wu, Huimin, et al.
Pubblicazione: (2025)
Semantic Gaussians: Open-Vocabulary Scene Understanding with 3D Gaussian Splatting
di: Guo, Jun, et al.
Pubblicazione: (2024)
di: Guo, Jun, et al.
Pubblicazione: (2024)
VideoAgent2: Enhancing the LLM-Based Agent System for Long-Form Video Understanding by Uncertainty-Aware CoT
di: Zhi, Zhuo, et al.
Pubblicazione: (2025)
di: Zhi, Zhuo, et al.
Pubblicazione: (2025)
From Waveforms to Pixels: A Survey on Audio-Visual Segmentation
di: Li, Jia, et al.
Pubblicazione: (2025)
di: Li, Jia, et al.
Pubblicazione: (2025)
FlowDreamer: A RGB-D World Model with Flow-based Motion Representations for Robot Manipulation
di: Guo, Jun, et al.
Pubblicazione: (2025)
di: Guo, Jun, et al.
Pubblicazione: (2025)
Benchmarking and Improving GUI Agents in High-Dynamic Environments
di: Liu, Enqi, et al.
Pubblicazione: (2026)
di: Liu, Enqi, et al.
Pubblicazione: (2026)
From Objects to Anywhere: A Holistic Benchmark for Multi-level Visual Grounding in 3D Scenes
di: Wang, Tianxu, et al.
Pubblicazione: (2025)
di: Wang, Tianxu, et al.
Pubblicazione: (2025)
Seeing Through Fog: Towards Fog-Invariant Action Recognition
di: Liu, Enqi, et al.
Pubblicazione: (2026)
di: Liu, Enqi, et al.
Pubblicazione: (2026)
STVG-R1: Incentivizing Instance-Level Reasoning and Grounding in Videos via Reinforcement Learning
di: Zhang, Xiaowen, et al.
Pubblicazione: (2026)
di: Zhang, Xiaowen, et al.
Pubblicazione: (2026)
SceneDreamer360: Text-Driven 3D-Consistent Scene Generation with Panoramic Gaussian Splatting
di: Li, Wenrui, et al.
Pubblicazione: (2024)
di: Li, Wenrui, et al.
Pubblicazione: (2024)
Unifying 3D Vision-Language Understanding via Promptable Queries
di: Zhu, Ziyu, et al.
Pubblicazione: (2024)
di: Zhu, Ziyu, et al.
Pubblicazione: (2024)
Building-PCC: Building Point Cloud Completion Benchmarks
di: Gao, Weixiao, et al.
Pubblicazione: (2024)
di: Gao, Weixiao, et al.
Pubblicazione: (2024)
CT-Agent: A Multimodal-LLM Agent for 3D CT Radiology Question Answering
di: Mao, Yuren, et al.
Pubblicazione: (2025)
di: Mao, Yuren, et al.
Pubblicazione: (2025)
SignLLM: Sign Language Production Large Language Models
di: Fang, Sen, et al.
Pubblicazione: (2024)
di: Fang, Sen, et al.
Pubblicazione: (2024)
GUIDE: Resolving Domain Bias in GUI Agents through Real-Time Web Video Retrieval and Plug-and-Play Annotation
di: Xie, Rui, et al.
Pubblicazione: (2026)
di: Xie, Rui, et al.
Pubblicazione: (2026)
LEO-VL: Efficient Scene Representation for Scalable 3D Vision-Language Learning
di: Huang, Jiangyong, et al.
Pubblicazione: (2025)
di: Huang, Jiangyong, et al.
Pubblicazione: (2025)
UltraEdit: Instruction-based Fine-Grained Image Editing at Scale
di: Zhao, Haozhe, et al.
Pubblicazione: (2024)
di: Zhao, Haozhe, et al.
Pubblicazione: (2024)
Mirror-3DGS: Incorporating Mirror Reflections into 3D Gaussian Splatting
di: Meng, Jiarui, et al.
Pubblicazione: (2024)
di: Meng, Jiarui, et al.
Pubblicazione: (2024)
A Multi-Agent System for Building-Age Cohort Mapping to Support Urban Energy Planning
di: Thota, Kundan, et al.
Pubblicazione: (2026)
di: Thota, Kundan, et al.
Pubblicazione: (2026)
Efficiently Leveraging Linguistic Priors for Scene Text Spotting
di: Nguyen, Nguyen, et al.
Pubblicazione: (2024)
di: Nguyen, Nguyen, et al.
Pubblicazione: (2024)
InternAgent: When Agent Becomes the Scientist -- Building Closed-Loop System from Hypothesis to Verification
di: InternAgent Team, et al.
Pubblicazione: (2025)
di: InternAgent Team, et al.
Pubblicazione: (2025)
Learning to Incorporate Texture Saliency Adaptive Attention to Image Cartoonization
di: Gao, Xiang, et al.
Pubblicazione: (2022)
di: Gao, Xiang, et al.
Pubblicazione: (2022)
Task-oriented Sequential Grounding and Navigation in 3D Scenes
di: Zhang, Zhuofan, et al.
Pubblicazione: (2024)
di: Zhang, Zhuofan, et al.
Pubblicazione: (2024)
An LLM-LVLM Driven Agent for Iterative and Fine-Grained Image Editing
di: Liang, Zihan, et al.
Pubblicazione: (2025)
di: Liang, Zihan, et al.
Pubblicazione: (2025)
Synergistic Perception and Generative Recomposition: A Multi-Agent Orchestration for Expert-Level Building Inspection
di: Zhong, Hui, et al.
Pubblicazione: (2026)
di: Zhong, Hui, et al.
Pubblicazione: (2026)
LossAgent: Towards Any Optimization Objectives for Image Processing with LLM Agents
di: Li, Bingchen, et al.
Pubblicazione: (2024)
di: Li, Bingchen, et al.
Pubblicazione: (2024)
Incorporating Scene Context and Semantic Labels for Enhanced Group-level Emotion Recognition
di: Zhu, Qing, et al.
Pubblicazione: (2025)
di: Zhu, Qing, et al.
Pubblicazione: (2025)
Incorporating Test-Time Optimization into Training with Dual Networks for Human Mesh Recovery
di: Nie, Yongwei, et al.
Pubblicazione: (2024)
di: Nie, Yongwei, et al.
Pubblicazione: (2024)
M$^2$: Dual-Memory Augmentation for Long-Horizon Web Agents via Trajectory Summarization and Insight Retrieval
di: Yan, Dawei, et al.
Pubblicazione: (2026)
di: Yan, Dawei, et al.
Pubblicazione: (2026)
Enhancing Multi-Agent Systems via Reinforcement Learning with LLM-based Planner and Graph-based Policy
di: Jia, Ziqi, et al.
Pubblicazione: (2025)
di: Jia, Ziqi, et al.
Pubblicazione: (2025)
FreSca: Scaling in Frequency Space Enhances Diffusion Models
di: Huang, Chao, et al.
Pubblicazione: (2025)
di: Huang, Chao, et al.
Pubblicazione: (2025)
An Embodied Generalist Agent in 3D World
di: Huang, Jiangyong, et al.
Pubblicazione: (2023)
di: Huang, Jiangyong, et al.
Pubblicazione: (2023)
UI-Venus Technical Report: Building High-performance UI Agents with RFT
di: Gu, Zhangxuan, et al.
Pubblicazione: (2025)
di: Gu, Zhangxuan, et al.
Pubblicazione: (2025)
Documenti analoghi
-
Iterative Tool Usage Exploration for Multimodal Agents via Step-wise Preference Tuning
di: Li, Pengxiang, et al.
Pubblicazione: (2025) -
VideoAgent: A Memory-augmented Multimodal Agent for Video Understanding
di: Fan, Yue, et al.
Pubblicazione: (2024) -
Multi-modal Agent Tuning: Building a VLM-Driven Agent for Efficient Tool Usage
di: Gao, Zhi, et al.
Pubblicazione: (2024) -
TongUI: Internet-Scale Trajectories from Multimodal Web Tutorials for Generalized GUI Agents
di: Zhang, Bofei, et al.
Pubblicazione: (2025) -
MILR: Improving Multimodal Image Generation via Test-Time Latent Reasoning
di: Mi, Yapeng, et al.
Pubblicazione: (2025)