PaperBanana: Automating Academic Illustration for AI Scientists
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Zhu, Dawei, Meng, Rui, Song, Yale, Wei, Xiyu, Li, Sujian, Pfister, Tomas, Yoon, Jinsung |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
DocLens : A Tool-Augmented Multi-Agent Framework for Long Visual Document Understanding
von: Zhu, Dawei, et al.
Veröffentlicht: (2025)
von: Zhu, Dawei, et al.
Veröffentlicht: (2025)
PaperOrchestra: A Multi-Agent Framework for Automated AI Research Paper Writing
von: Song, Yiwen, et al.
Veröffentlicht: (2026)
von: Song, Yiwen, et al.
Veröffentlicht: (2026)
VQQA: An Agentic Approach for Video Evaluation and Quality Improvement
von: Song, Yiwen, et al.
Veröffentlicht: (2026)
von: Song, Yiwen, et al.
Veröffentlicht: (2026)
ScientistOne: Towards Human-Level Autonomous Research via Chain-of-Evidence
von: Meng, Rui, et al.
Veröffentlicht: (2026)
von: Meng, Rui, et al.
Veröffentlicht: (2026)
Relevance-aware Multi-context Contrastive Decoding for Retrieval-augmented Visual Question Answering
von: Kim, Jongha, et al.
Veröffentlicht: (2026)
von: Kim, Jongha, et al.
Veröffentlicht: (2026)
Jr. AI Scientist and Its Risk Report: Autonomous Scientific Exploration from a Baseline Paper
von: Miyai, Atsuyuki, et al.
Veröffentlicht: (2025)
von: Miyai, Atsuyuki, et al.
Veröffentlicht: (2025)
VITED: Video Temporal Evidence Distillation
von: Lu, Yujie, et al.
Veröffentlicht: (2025)
von: Lu, Yujie, et al.
Veröffentlicht: (2025)
LLMs Behind the Scenes: Enabling Narrative Scene Illustration
von: Roemmele, Melissa, et al.
Veröffentlicht: (2025)
von: Roemmele, Melissa, et al.
Veröffentlicht: (2025)
CROME: Cross-Modal Adapters for Efficient Multimodal LLM
von: Ebrahimi, Sayna, et al.
Veröffentlicht: (2024)
von: Ebrahimi, Sayna, et al.
Veröffentlicht: (2024)
A$^2$RD: Agentic Autoregressive Diffusion for Long Video Consistency
von: Long, Do Xuan, et al.
Veröffentlicht: (2026)
von: Long, Do Xuan, et al.
Veröffentlicht: (2026)
Paper2Poster: Towards Multimodal Poster Automation from Scientific Papers
von: Pang, Wei, et al.
Veröffentlicht: (2025)
von: Pang, Wei, et al.
Veröffentlicht: (2025)
Matryoshka-Adaptor: Unsupervised and Supervised Tuning for Smaller Embedding Dimensions
von: Yoon, Jinsung, et al.
Veröffentlicht: (2024)
von: Yoon, Jinsung, et al.
Veröffentlicht: (2024)
VL-RewardBench: A Challenging Benchmark for Vision-Language Generative Reward Models
von: Li, Lei, et al.
Veröffentlicht: (2024)
von: Li, Lei, et al.
Veröffentlicht: (2024)
Pico-Banana-400K: A Large-Scale Dataset for Text-Guided Image Editing
von: Qian, Yusu, et al.
Veröffentlicht: (2025)
von: Qian, Yusu, et al.
Veröffentlicht: (2025)
To Preserve or To Compress: An In-Depth Study of Connector Selection in Multimodal Large Language Models
von: Lin, Junyan, et al.
Veröffentlicht: (2024)
von: Lin, Junyan, et al.
Veröffentlicht: (2024)
Transformers and Language Models in Form Understanding: A Comprehensive Review of Scanned Document Analysis
von: Abdallah, Abdelrahman, et al.
Veröffentlicht: (2024)
von: Abdallah, Abdelrahman, et al.
Veröffentlicht: (2024)
AutoFigure: Generating and Refining Publication-Ready Scientific Illustrations
von: Zhu, Minjun, et al.
Veröffentlicht: (2026)
von: Zhu, Minjun, et al.
Veröffentlicht: (2026)
SciGA: A Comprehensive Dataset for Designing Graphical Abstracts in Academic Papers
von: Kawada, Takuro, et al.
Veröffentlicht: (2025)
von: Kawada, Takuro, et al.
Veröffentlicht: (2025)
PolicyBank: Evolving Policy Understanding for LLM Agents
von: Choi, Jihye, et al.
Veröffentlicht: (2026)
von: Choi, Jihye, et al.
Veröffentlicht: (2026)
ATLAS: Constraints-Aware Multi-Agent Collaboration for Real-World Travel Planning
von: Choi, Jihye, et al.
Veröffentlicht: (2025)
von: Choi, Jihye, et al.
Veröffentlicht: (2025)
Paper2Web: Let's Make Your Paper Alive!
von: Chen, Yuhang, et al.
Veröffentlicht: (2025)
von: Chen, Yuhang, et al.
Veröffentlicht: (2025)
Assessing News Thumbnail Representativeness: Counterfactual text can enhance the cross-modal matching ability
von: Yoon, Yejun, et al.
Veröffentlicht: (2024)
von: Yoon, Yejun, et al.
Veröffentlicht: (2024)
InternAgent: When Agent Becomes the Scientist -- Building Closed-Loop System from Hypothesis to Verification
von: InternAgent Team, et al.
Veröffentlicht: (2025)
von: InternAgent Team, et al.
Veröffentlicht: (2025)
AIBench: Evaluating Visual-Logical Consistency in Academic Illustration Generation
von: Liao, Zhaohe, et al.
Veröffentlicht: (2026)
von: Liao, Zhaohe, et al.
Veröffentlicht: (2026)
Video2GUI: Synthesizing Large-Scale Interaction Trajectories for Generalized GUI Agent Pretraining
von: Xiong, Weimin, et al.
Veröffentlicht: (2026)
von: Xiong, Weimin, et al.
Veröffentlicht: (2026)
Task-Specific Directions: Definition, Exploration, and Utilization in Parameter Efficient Fine-Tuning
von: Si, Chongjie, et al.
Veröffentlicht: (2024)
von: Si, Chongjie, et al.
Veröffentlicht: (2024)
MMLongBench: Benchmarking Long-Context Vision-Language Models Effectively and Thoroughly
von: Wang, Zhaowei, et al.
Veröffentlicht: (2025)
von: Wang, Zhaowei, et al.
Veröffentlicht: (2025)
Know "No" Better: A Data-Driven Approach for Enhancing Negation Awareness in CLIP
von: Park, Junsung, et al.
Veröffentlicht: (2025)
von: Park, Junsung, et al.
Veröffentlicht: (2025)
Multi-Modal Explainable Medical AI Assistant for Trustworthy Human-AI Collaboration
von: Yang, Honglong, et al.
Veröffentlicht: (2025)
von: Yang, Honglong, et al.
Veröffentlicht: (2025)
Can Video LLMs Refuse to Answer? Alignment for Answerability in Video Large Language Models
von: Yoon, Eunseop, et al.
Veröffentlicht: (2025)
von: Yoon, Eunseop, et al.
Veröffentlicht: (2025)
Hierarchy-Aware Multimodal Unlearning for Medical AI
von: Wu, Fengli, et al.
Veröffentlicht: (2025)
von: Wu, Fengli, et al.
Veröffentlicht: (2025)
Open-Source Image Editing Models Are Zero-Shot Vision Learners
von: Liu, Wei, et al.
Veröffentlicht: (2026)
von: Liu, Wei, et al.
Veröffentlicht: (2026)
E$^2$AT: Multimodal Jailbreak Defense via Dynamic Joint Optimization for Multimodal Large Language Models
von: Lu, Liming, et al.
Veröffentlicht: (2025)
von: Lu, Liming, et al.
Veröffentlicht: (2025)
Are Vision LLMs Road-Ready? A Comprehensive Benchmark for Safety-Critical Driving Video Understanding
von: Zeng, Tong, et al.
Veröffentlicht: (2025)
von: Zeng, Tong, et al.
Veröffentlicht: (2025)
Octopus v3: Technical Report for On-device Sub-billion Multimodal AI Agent
von: Chen, Wei, et al.
Veröffentlicht: (2024)
von: Chen, Wei, et al.
Veröffentlicht: (2024)
Towards Zero-Shot Annotation of the Built Environment with Vision-Language Models (Vision Paper)
von: Han, Bin, et al.
Veröffentlicht: (2024)
von: Han, Bin, et al.
Veröffentlicht: (2024)
Modularized Networks for Few-shot Hateful Meme Detection
von: Cao, Rui, et al.
Veröffentlicht: (2024)
von: Cao, Rui, et al.
Veröffentlicht: (2024)
VLMs-in-the-Wild: Bridging the Gap Between Academic Benchmarks and Enterprise Reality
von: Bandraupalli, Srihari, et al.
Veröffentlicht: (2025)
von: Bandraupalli, Srihari, et al.
Veröffentlicht: (2025)
MultiChartQA: Benchmarking Vision-Language Models on Multi-Chart Problems
von: Zhu, Zifeng, et al.
Veröffentlicht: (2024)
von: Zhu, Zifeng, et al.
Veröffentlicht: (2024)
TUMIX: Multi-Agent Test-Time Scaling with Tool-Use Mixture
von: Chen, Yongchao, et al.
Veröffentlicht: (2025)
von: Chen, Yongchao, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
DocLens : A Tool-Augmented Multi-Agent Framework for Long Visual Document Understanding
von: Zhu, Dawei, et al.
Veröffentlicht: (2025) -
PaperOrchestra: A Multi-Agent Framework for Automated AI Research Paper Writing
von: Song, Yiwen, et al.
Veröffentlicht: (2026) -
VQQA: An Agentic Approach for Video Evaluation and Quality Improvement
von: Song, Yiwen, et al.
Veröffentlicht: (2026) -
ScientistOne: Towards Human-Level Autonomous Research via Chain-of-Evidence
von: Meng, Rui, et al.
Veröffentlicht: (2026) -
Relevance-aware Multi-context Contrastive Decoding for Retrieval-augmented Visual Question Answering
von: Kim, Jongha, et al.
Veröffentlicht: (2026)