A Unified Agentic Framework for Evaluating Conditional Image Generation
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Wang, Jifang, Yang, Xue, Wang, Longyue, Xu, Zhenran, Wang, Yiyu, Wang, Yaowei, Luo, Weihua, Zhang, Kaifu, Hu, Baotian, Zhang, Min |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
ComfyUI-R1: Exploring Reasoning Models for Workflow Generation
von: Xu, Zhenran, et al.
Veröffentlicht: (2025)
von: Xu, Zhenran, et al.
Veröffentlicht: (2025)
ComfyUI-Copilot: An Intelligent Assistant for Automated Workflow Development
von: Xu, Zhenran, et al.
Veröffentlicht: (2025)
von: Xu, Zhenran, et al.
Veröffentlicht: (2025)
FilmAgent: A Multi-Agent Framework for End-to-End Film Automation in Virtual 3D Spaces
von: Xu, Zhenran, et al.
Veröffentlicht: (2025)
von: Xu, Zhenran, et al.
Veröffentlicht: (2025)
Generative Multimodal Entity Linking
von: Shi, Senbao, et al.
Veröffentlicht: (2023)
von: Shi, Senbao, et al.
Veröffentlicht: (2023)
Towards Lightweight, Adaptive and Attribute-Aware Multi-Aspect Controllable Text Generation with Large Language Models
von: Zhu, Chenyu, et al.
Veröffentlicht: (2025)
von: Zhu, Chenyu, et al.
Veröffentlicht: (2025)
DeepWideSearch: Benchmarking Depth and Width in Agentic Information Seeking
von: Lan, Tian, et al.
Veröffentlicht: (2025)
von: Lan, Tian, et al.
Veröffentlicht: (2025)
(Perhaps) Beyond Human Translation: Harnessing Multi-Agent Collaboration for Translating Ultra-Long Literary Texts
von: Wu, Minghao, et al.
Veröffentlicht: (2024)
von: Wu, Minghao, et al.
Veröffentlicht: (2024)
Picking the Cream of the Crop: Visual-Centric Data Selection with Collaborative Agents
von: Liu, Zhenyu, et al.
Veröffentlicht: (2025)
von: Liu, Zhenyu, et al.
Veröffentlicht: (2025)
New Trends for Modern Machine Translation with Large Reasoning Models
von: Liu, Sinuo, et al.
Veröffentlicht: (2025)
von: Liu, Sinuo, et al.
Veröffentlicht: (2025)
MSVBench: Towards Human-Level Evaluation of Multi-Shot Video Generation
von: Shi, Haoyuan, et al.
Veröffentlicht: (2026)
von: Shi, Haoyuan, et al.
Veröffentlicht: (2026)
A Comprehensive Evaluation of GPT-4V on Knowledge-Intensive Visual Question Answering
von: Li, Yunxin, et al.
Veröffentlicht: (2023)
von: Li, Yunxin, et al.
Veröffentlicht: (2023)
Uni-MoE: Scaling Unified Multimodal LLMs with Mixture of Experts
von: Li, Yunxin, et al.
Veröffentlicht: (2024)
von: Li, Yunxin, et al.
Veröffentlicht: (2024)
VisionGraph: Leveraging Large Multimodal Models for Graph Theory Problems in Visual Context
von: Li, Yunxin, et al.
Veröffentlicht: (2024)
von: Li, Yunxin, et al.
Veröffentlicht: (2024)
Challenging Multilingual LLMs: A New Taxonomy and Benchmark for Unraveling Hallucination in Translation
von: Wu, Xinwei, et al.
Veröffentlicht: (2025)
von: Wu, Xinwei, et al.
Veröffentlicht: (2025)
Perception, Reason, Think, and Plan: A Survey on Large Multimodal Reasoning Models
von: Li, Yunxin, et al.
Veröffentlicht: (2025)
von: Li, Yunxin, et al.
Veröffentlicht: (2025)
WindowsWorld: A Process-Centric Benchmark of Autonomous GUI Agents in Professional Cross-Application Environments
von: Li, Jinchao, et al.
Veröffentlicht: (2026)
von: Li, Jinchao, et al.
Veröffentlicht: (2026)
Rethinking Multilingual Vision-Language Translation: Dataset, Evaluation, and Adaptation
von: Wang, Xintong, et al.
Veröffentlicht: (2025)
von: Wang, Xintong, et al.
Veröffentlicht: (2025)
LayAlign: Enhancing Multilingual Reasoning in Large Language Models via Layer-Wise Adaptive Fusion and Alignment Strategy
von: Ruan, Zhiwen, et al.
Veröffentlicht: (2025)
von: Ruan, Zhiwen, et al.
Veröffentlicht: (2025)
Marco-o1: Towards Open Reasoning Models for Open-Ended Solutions
von: Zhao, Yu, et al.
Veröffentlicht: (2024)
von: Zhao, Yu, et al.
Veröffentlicht: (2024)
VideoVista: A Versatile Benchmark for Video Understanding and Reasoning
von: Li, Yunxin, et al.
Veröffentlicht: (2024)
von: Li, Yunxin, et al.
Veröffentlicht: (2024)
VideoVista-CulturalLingo: 360$^\circ$ Horizons-Bridging Cultures, Languages, and Domains in Video Comprehension
von: Chen, Xinyu, et al.
Veröffentlicht: (2025)
von: Chen, Xinyu, et al.
Veröffentlicht: (2025)
Anim-Director: A Large Multimodal Model Powered Agent for Controllable Animation Video Generation
von: Li, Yunxin, et al.
Veröffentlicht: (2024)
von: Li, Yunxin, et al.
Veröffentlicht: (2024)
Beyond Single-Reward: Multi-Pair, Multi-Perspective Preference Optimization for Machine Translation
von: Wang, Hao, et al.
Veröffentlicht: (2025)
von: Wang, Hao, et al.
Veröffentlicht: (2025)
The Bitter Lesson Learned from 2,000+ Multilingual Benchmarks
von: Wu, Minghao, et al.
Veröffentlicht: (2025)
von: Wu, Minghao, et al.
Veröffentlicht: (2025)
Structured Episodic Event Memory
von: Lu, Zhengxuan, et al.
Veröffentlicht: (2026)
von: Lu, Zhengxuan, et al.
Veröffentlicht: (2026)
VerIPO: Cultivating Long Reasoning in Video-LLMs via Verifier-Gudied Iterative Policy Optimization
von: Li, Yunxin, et al.
Veröffentlicht: (2025)
von: Li, Yunxin, et al.
Veröffentlicht: (2025)
ToolOmni: Enabling Open-World Tool Use via Agentic learning with Proactive Retrieval and Grounded Execution
von: Huang, Shouzheng, et al.
Veröffentlicht: (2026)
von: Huang, Shouzheng, et al.
Veröffentlicht: (2026)
Finding the Translation Switch: Discovering and Exploiting the Task-Initiation Features in LLMs
von: Wu, Xinwei, et al.
Veröffentlicht: (2026)
von: Wu, Xinwei, et al.
Veröffentlicht: (2026)
UMEM: Unified Memory Extraction and Management Framework for Generalizable Memory
von: Ye, Yongshi, et al.
Veröffentlicht: (2026)
von: Ye, Yongshi, et al.
Veröffentlicht: (2026)
Building Autonomous GUI Navigation via Agentic-Q Estimation and Step-Wise Policy Optimization
von: Wang, Yibo, et al.
Veröffentlicht: (2026)
von: Wang, Yibo, et al.
Veröffentlicht: (2026)
Medico: Towards Hallucination Detection and Correction with Multi-source Evidence Fusion
von: Zhao, Xinping, et al.
Veröffentlicht: (2024)
von: Zhao, Xinping, et al.
Veröffentlicht: (2024)
HSCodeComp: A Realistic and Expert-level Benchmark for Deep Search Agents in Hierarchical Rule Application
von: Yang, Yiqian, et al.
Veröffentlicht: (2025)
von: Yang, Yiqian, et al.
Veröffentlicht: (2025)
Table-as-Search: Formulate Long-Horizon Agentic Information Seeking as Table Completion
von: Lan, Tian, et al.
Veröffentlicht: (2026)
von: Lan, Tian, et al.
Veröffentlicht: (2026)
Marco-LLM: Bridging Languages via Massive Multilingual Training for Cross-Lingual Enhancement
von: Ming, Lingfeng, et al.
Veröffentlicht: (2024)
von: Ming, Lingfeng, et al.
Veröffentlicht: (2024)
Can LLMs Track Their Output Length? A Dynamic Feedback Mechanism for Precise Length Regulation
von: Xiao, Meiman, et al.
Veröffentlicht: (2026)
von: Xiao, Meiman, et al.
Veröffentlicht: (2026)
VIDA: A dataset for Visually Dependent Ambiguity in Multimodal Machine Translation
von: Pan, Jingheng, et al.
Veröffentlicht: (2026)
von: Pan, Jingheng, et al.
Veröffentlicht: (2026)
Less Languages, Less Tokens: An Efficient Unified Logic Cross-lingual Chain-of-Thought Reasoning Framework
von: Zhang, Chenyuan, et al.
Veröffentlicht: (2026)
von: Zhang, Chenyuan, et al.
Veröffentlicht: (2026)
Triplets Better Than Pairs: Towards Stable and Effective Self-Play Fine-Tuning for LLMs
von: Wang, Yibo, et al.
Veröffentlicht: (2026)
von: Wang, Yibo, et al.
Veröffentlicht: (2026)
Marco-Voice Technical Report
von: Tian, Fengping, et al.
Veröffentlicht: (2025)
von: Tian, Fengping, et al.
Veröffentlicht: (2025)
A Multimodal In-Context Tuning Approach for E-Commerce Product Description Generation
von: Li, Yunxin, et al.
Veröffentlicht: (2024)
von: Li, Yunxin, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
ComfyUI-R1: Exploring Reasoning Models for Workflow Generation
von: Xu, Zhenran, et al.
Veröffentlicht: (2025) -
ComfyUI-Copilot: An Intelligent Assistant for Automated Workflow Development
von: Xu, Zhenran, et al.
Veröffentlicht: (2025) -
FilmAgent: A Multi-Agent Framework for End-to-End Film Automation in Virtual 3D Spaces
von: Xu, Zhenran, et al.
Veröffentlicht: (2025) -
Generative Multimodal Entity Linking
von: Shi, Senbao, et al.
Veröffentlicht: (2023) -
Towards Lightweight, Adaptive and Attribute-Aware Multi-Aspect Controllable Text Generation with Large Language Models
von: Zhu, Chenyu, et al.
Veröffentlicht: (2025)