Generating Coherent Sequences of Visual Illustrations for Real-World Manual Tasks
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Bordalo, João, Ramos, Vasco, Valério, Rodrigo, Glória-Silva, Diogo, Bitton, Yonatan, Yarom, Michal, Szpektor, Idan, Magalhaes, Joao |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Contrastive Sequential-Diffusion Learning: Non-linear and Multi-Scene Instructional Video Synthesis
von: Ramos, Vasco, et al.
Veröffentlicht: (2024)
von: Ramos, Vasco, et al.
Veröffentlicht: (2024)
Latent Beam Diffusion Models for Generating Visual Sequences
von: Fernandes, Guilherme, et al.
Veröffentlicht: (2025)
von: Fernandes, Guilherme, et al.
Veröffentlicht: (2025)
TALC: Time-Aligned Captions for Multi-Scene Text-to-Video Generation
von: Bansal, Hritik, et al.
Veröffentlicht: (2024)
von: Bansal, Hritik, et al.
Veröffentlicht: (2024)
Beyond the Noise: Aligning Prompts with Latent Representations in Diffusion Models
von: Ramos, Vasco, et al.
Veröffentlicht: (2025)
von: Ramos, Vasco, et al.
Veröffentlicht: (2025)
TWIZ-v2: The Wizard of Multimodal Conversational-Stimulus
von: Ferreira, Rafael, et al.
Veröffentlicht: (2023)
von: Ferreira, Rafael, et al.
Veröffentlicht: (2023)
Bridging the Visual Gap: Fine-Tuning Multimodal Models with Knowledge-Adapted Captions
von: Yanuka, Moran, et al.
Veröffentlicht: (2024)
von: Yanuka, Moran, et al.
Veröffentlicht: (2024)
Visual Riddles: a Commonsense and World Knowledge Challenge for Large Vision and Language Models
von: Bitton-Guetta, Nitzan, et al.
Veröffentlicht: (2024)
von: Bitton-Guetta, Nitzan, et al.
Veröffentlicht: (2024)
Mismatch Quest: Visual and Textual Feedback for Image-Text Misalignment
von: Gordon, Brian, et al.
Veröffentlicht: (2023)
von: Gordon, Brian, et al.
Veröffentlicht: (2023)
Video-STaR: Self-Training Enables Video Instruction Tuning with Any Supervision
von: Zohar, Orr, et al.
Veröffentlicht: (2024)
von: Zohar, Orr, et al.
Veröffentlicht: (2024)
Seeing Isn't Knowing: Do VLMs Know When Not to Answer Spatial Questions (and Why)?
von: Zhang, Yue, et al.
Veröffentlicht: (2026)
von: Zhang, Yue, et al.
Veröffentlicht: (2026)
RefVNLI: Towards Scalable Evaluation of Subject-driven Text-to-image Generation
von: Slobodkin, Aviv, et al.
Veröffentlicht: (2025)
von: Slobodkin, Aviv, et al.
Veröffentlicht: (2025)
Unblocking Fine-Grained Evaluation of Detailed Captions: An Explaining AutoRater and Critic-and-Revise Pipeline
von: Gordon, Brian, et al.
Veröffentlicht: (2025)
von: Gordon, Brian, et al.
Veröffentlicht: (2025)
Distinguishing Ignorance from Error in LLM Hallucinations
von: Simhi, Adi, et al.
Veröffentlicht: (2024)
von: Simhi, Adi, et al.
Veröffentlicht: (2024)
Constructing Benchmarks and Interventions for Combating Hallucinations in LLMs
von: Simhi, Adi, et al.
Veröffentlicht: (2024)
von: Simhi, Adi, et al.
Veröffentlicht: (2024)
Error-Driven Scene Editing for 3D Grounding in Large Language Models
von: Zhang, Yue, et al.
Veröffentlicht: (2025)
von: Zhang, Yue, et al.
Veröffentlicht: (2025)
Show and Guide: Instructional-Plan Grounded Vision and Language Model
von: Glória-Silva, Diogo, et al.
Veröffentlicht: (2024)
von: Glória-Silva, Diogo, et al.
Veröffentlicht: (2024)
VideoPhy: Evaluating Physical Commonsense for Video Generation
von: Bansal, Hritik, et al.
Veröffentlicht: (2024)
von: Bansal, Hritik, et al.
Veröffentlicht: (2024)
Plan-Grounded Large Language Models for Dual Goal Conversational Settings
von: Glória-Silva, Diogo, et al.
Veröffentlicht: (2024)
von: Glória-Silva, Diogo, et al.
Veröffentlicht: (2024)
ManagerBench: Evaluating the Safety-Pragmatism Trade-off in Autonomous LLMs
von: Simhi, Adi, et al.
Veröffentlicht: (2025)
von: Simhi, Adi, et al.
Veröffentlicht: (2025)
3DLLM-Mem: Long-Term Spatial-Temporal Memory for Embodied 3D Large Language Model
von: Hu, Wenbo, et al.
Veröffentlicht: (2025)
von: Hu, Wenbo, et al.
Veröffentlicht: (2025)
LLMs Know More Than They Show: On the Intrinsic Representation of LLM Hallucinations
von: Orgad, Hadas, et al.
Veröffentlicht: (2024)
von: Orgad, Hadas, et al.
Veröffentlicht: (2024)
Microbiological water quality in urban coastal beaches: the influence of water dynamics and optimization of the sampling strategy
von: Bordalo, A
Veröffentlicht: (2001)
von: Bordalo, A
Veröffentlicht: (2001)
VIGiA: Instructional Video Guidance via Dialogue Reasoning and Retrieval
von: Glória-Silva, Diogo, et al.
Veröffentlicht: (2026)
von: Glória-Silva, Diogo, et al.
Veröffentlicht: (2026)
ParallelPARC: A Scalable Pipeline for Generating Natural-Language Analogies
von: Sultan, Oren, et al.
Veröffentlicht: (2024)
von: Sultan, Oren, et al.
Veröffentlicht: (2024)
An Efficient and Effective Encoder Model for Vision and Language Tasks in the Remote Sensing Domain
von: Silva, João Daniel, et al.
Veröffentlicht: (2025)
von: Silva, João Daniel, et al.
Veröffentlicht: (2025)
ALBA: A European Portuguese Benchmark for Evaluating Language and Linguistic Dimensions in Generative LLMs
von: Vieira, Inês, et al.
Veröffentlicht: (2026)
von: Vieira, Inês, et al.
Veröffentlicht: (2026)
MoralCLIP: Contrastive Alignment of Vision-and-Language Representations with Moral Foundations Theory
von: Condez, Ana Carolina, et al.
Veröffentlicht: (2025)
von: Condez, Ana Carolina, et al.
Veröffentlicht: (2025)
Security Steerability is All You Need
von: Hazan, Itay, et al.
Veröffentlicht: (2025)
von: Hazan, Itay, et al.
Veröffentlicht: (2025)
RedTWIZ: Diverse LLM Red Teaming via Adaptive Attack Planning
von: Horal, Artur, et al.
Veröffentlicht: (2025)
von: Horal, Artur, et al.
Veröffentlicht: (2025)
Inside-Out: Hidden Factual Knowledge in LLMs
von: Gekhman, Zorik, et al.
Veröffentlicht: (2025)
von: Gekhman, Zorik, et al.
Veröffentlicht: (2025)
StreetView-Waste: A Multi-Task Dataset for Urban Waste Management
von: Paulo, Diogo J., et al.
Veröffentlicht: (2025)
von: Paulo, Diogo J., et al.
Veröffentlicht: (2025)
Hands-off Image Editing: Language-guided Editing without any Task-specific Labeling, Masking or even Training
von: Santos, Rodrigo, et al.
Veröffentlicht: (2025)
von: Santos, Rodrigo, et al.
Veröffentlicht: (2025)
EditInspector: A Benchmark for Evaluation of Text-Guided Image Edits
von: Yosef, Ron, et al.
Veröffentlicht: (2025)
von: Yosef, Ron, et al.
Veröffentlicht: (2025)
Are LLMs Better than Reported? Detecting Label Errors and Mitigating Their Effect on Model Performance
von: Nahum, Omer, et al.
Veröffentlicht: (2024)
von: Nahum, Omer, et al.
Veröffentlicht: (2024)
Beneath the Surface of Consistency: Exploring Cross-lingual Knowledge Representation Sharing in LLMs
von: Ifergan, Maxim, et al.
Veröffentlicht: (2024)
von: Ifergan, Maxim, et al.
Veröffentlicht: (2024)
Illustrated Manual of Pediatric Dermatology
von: Mallory, Susan, et al.
Veröffentlicht: (2020)
von: Mallory, Susan, et al.
Veröffentlicht: (2020)
Multi-trait User Simulation with Adaptive Decoding for Conversational Task Assistants
von: Ferreira, Rafael, et al.
Veröffentlicht: (2024)
von: Ferreira, Rafael, et al.
Veröffentlicht: (2024)
A gestão dos recursos hídricos a luz da ecologia política: um debate sobre o controle público versus o controle privado da água no Brasil
von: Carlos Alexandre Leão Bordalo
Veröffentlicht: (2008)
von: Carlos Alexandre Leão Bordalo
Veröffentlicht: (2008)
Simple grammar bisimilarity, with an application to session type equivalence
von: Poças, Diogo, et al.
Veröffentlicht: (2024)
von: Poças, Diogo, et al.
Veröffentlicht: (2024)
FineVAU: A Novel Human-Aligned Benchmark for Fine-Grained Video Anomaly Understanding
von: Pereira, João, et al.
Veröffentlicht: (2026)
von: Pereira, João, et al.
Veröffentlicht: (2026)
Ähnliche Einträge
-
Contrastive Sequential-Diffusion Learning: Non-linear and Multi-Scene Instructional Video Synthesis
von: Ramos, Vasco, et al.
Veröffentlicht: (2024) -
Latent Beam Diffusion Models for Generating Visual Sequences
von: Fernandes, Guilherme, et al.
Veröffentlicht: (2025) -
TALC: Time-Aligned Captions for Multi-Scene Text-to-Video Generation
von: Bansal, Hritik, et al.
Veröffentlicht: (2024) -
Beyond the Noise: Aligning Prompts with Latent Representations in Diffusion Models
von: Ramos, Vasco, et al.
Veröffentlicht: (2025) -
TWIZ-v2: The Wizard of Multimodal Conversational-Stimulus
von: Ferreira, Rafael, et al.
Veröffentlicht: (2023)