Text-to-Scene with Large Reasoning Models
Fuente:
arXiv
Saved in:
| Main Authors: | Berdoz, Frédéric, Lanzendörfer, Luca A., Tuninga, Nick, Wattenhofer, Roger |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
WorldSpeech: A Multilingual Speech Corpus from Around the World
by: Asonitis, Antonis, et al.
Published: (2026)
by: Asonitis, Antonis, et al.
Published: (2026)
Alignment-Aware Decoding
by: Berdoz, Frédéric, et al.
Published: (2025)
by: Berdoz, Frédéric, et al.
Published: (2025)
Virtual Fashion Photo-Shoots: Building a Large-Scale Garment-Lookbook Dataset
by: Hauri, Yannick, et al.
Published: (2025)
by: Hauri, Yannick, et al.
Published: (2025)
High-Fidelity Speech Enhancement via Discrete Audio Tokens
by: Lanzendörfer, Luca A., et al.
Published: (2025)
by: Lanzendörfer, Luca A., et al.
Published: (2025)
FLIP Reasoning Challenge
by: Plesner, Andreas, et al.
Published: (2025)
by: Plesner, Andreas, et al.
Published: (2025)
AEye: A Visualization Tool for Image Datasets
by: Grötschla, Florian, et al.
Published: (2024)
by: Grötschla, Florian, et al.
Published: (2024)
Efficient Bayesian Inference from Noisy Pairwise Comparisons
by: Aczel, Till, et al.
Published: (2025)
by: Aczel, Till, et al.
Published: (2025)
Reasoning Boosts Opinion Alignment in LLMs
by: Berdoz, Frédéric, et al.
Published: (2026)
by: Berdoz, Frédéric, et al.
Published: (2026)
The Unwinnable Arms Race of AI Image Detection
by: Aczel, Till, et al.
Published: (2025)
by: Aczel, Till, et al.
Published: (2025)
Conditional Hallucinations for Image Compression
by: Aczel, Till, et al.
Published: (2024)
by: Aczel, Till, et al.
Published: (2024)
From MNIST to ImageNet: Understanding the Scalability Boundaries of Differentiable Logic Gate Networks
by: Brändle, Sven, et al.
Published: (2025)
by: Brändle, Sven, et al.
Published: (2025)
SUPClust: Active Learning at the Boundaries
by: Ono, Yuta, et al.
Published: (2024)
by: Ono, Yuta, et al.
Published: (2024)
Bridging Diversity and Uncertainty in Active learning with Self-Supervised Pre-Training
by: Doucet, Paul, et al.
Published: (2024)
by: Doucet, Paul, et al.
Published: (2024)
What's in Common? Multimodal Models Hallucinate When Reasoning Across Scenes
by: Ross, Candace, et al.
Published: (2025)
by: Ross, Candace, et al.
Published: (2025)
SceneAlign: Aligning Multimodal Reasoning to Scene Graphs in Complex Visual Scenes
by: Wang, Chuhan, et al.
Published: (2026)
by: Wang, Chuhan, et al.
Published: (2026)
Keep It Real: Challenges in Attacking Compression-Based Adversarial Purification
by: Räber, Samuel, et al.
Published: (2025)
by: Räber, Samuel, et al.
Published: (2025)
ImageChain: Advancing Sequential Image-to-Text Reasoning in Multimodal Large Language Models
by: Villegas, Danae Sánchez, et al.
Published: (2025)
by: Villegas, Danae Sánchez, et al.
Published: (2025)
Ensemble Learning for Vietnamese Scene Text Spotting in Urban Environments
by: Nguyen, Hieu, et al.
Published: (2024)
by: Nguyen, Hieu, et al.
Published: (2024)
STORM: Spatio-Temporal Reconstruction Model for Large-Scale Outdoor Scenes
by: Yang, Jiawei, et al.
Published: (2024)
by: Yang, Jiawei, et al.
Published: (2024)
Steering Pretrained Drafters during Speculative Decoding
by: Berdoz, Frédéric, et al.
Published: (2025)
by: Berdoz, Frédéric, et al.
Published: (2025)
ReLaX: Reasoning with Latent Exploration for Large Reasoning Models
by: Zhang, Shimin, et al.
Published: (2025)
by: Zhang, Shimin, et al.
Published: (2025)
Generating Compositional Scenes via Text-to-image RGBA Instance Generation
by: Fontanella, Alessandro, et al.
Published: (2024)
by: Fontanella, Alessandro, et al.
Published: (2024)
Evaluating Numerical Reasoning in Text-to-Image Models
by: Kajić, Ivana, et al.
Published: (2024)
by: Kajić, Ivana, et al.
Published: (2024)
Can AI Agents Agree?
by: Berdoz, Frédéric, et al.
Published: (2026)
by: Berdoz, Frédéric, et al.
Published: (2026)
The Impact of Scaling Training Data on Adversarial Robustness
by: Zimmerli, Marco, et al.
Published: (2025)
by: Zimmerli, Marco, et al.
Published: (2025)
Detect in Any Scene: An Agentic Framework for Object Detection with Experience-Aware Reasoning
by: Zhang, Wenlun, et al.
Published: (2026)
by: Zhang, Wenlun, et al.
Published: (2026)
Lumos : Empowering Multimodal LLMs with Scene Text Recognition
by: Shenoy, Ashish, et al.
Published: (2024)
by: Shenoy, Ashish, et al.
Published: (2024)
Sum-of-Checks: Structured Reasoning for Surgical Safety with Large Vision-Language Models
by: You, Weiqiu, et al.
Published: (2026)
by: You, Weiqiu, et al.
Published: (2026)
GeoReasoner: Geo-localization with Reasoning in Street Views using a Large Vision-Language Model
by: Li, Ling, et al.
Published: (2024)
by: Li, Ling, et al.
Published: (2024)
ConTextual: Evaluating Context-Sensitive Text-Rich Visual Reasoning in Large Multimodal Models
by: Wadhawan, Rohan, et al.
Published: (2024)
by: Wadhawan, Rohan, et al.
Published: (2024)
Can Vision-Language Models See Squares? Text-Recognition Mediates Spatial Reasoning Across Three Model Families
by: Levental, Yuval
Published: (2026)
by: Levental, Yuval
Published: (2026)
LLM-Seg: Bridging Image Segmentation and Large Language Model Reasoning
by: Wang, Junchi, et al.
Published: (2024)
by: Wang, Junchi, et al.
Published: (2024)
Text-to-CAD Generation Through Infusing Visual Feedback in Large Language Models
by: Wang, Ruiyu, et al.
Published: (2025)
by: Wang, Ruiyu, et al.
Published: (2025)
LVT: Large-Scale Scene Reconstruction via Local View Transformers
by: Imtiaz, Tooba, et al.
Published: (2025)
by: Imtiaz, Tooba, et al.
Published: (2025)
Sora Detector: A Unified Hallucination Detection for Large Text-to-Video Models
by: Chu, Zhixuan, et al.
Published: (2024)
by: Chu, Zhixuan, et al.
Published: (2024)
Simple Vision-Language Math Reasoning via Rendered Text
by: Skripkin, Matvey, et al.
Published: (2025)
by: Skripkin, Matvey, et al.
Published: (2025)
VCR: A Task for Pixel-Level Complex Reasoning in Vision Language Models via Restoring Occluded Text
by: Zhang, Tianyu, et al.
Published: (2024)
by: Zhang, Tianyu, et al.
Published: (2024)
LaVR: Scene Latent Conditioned Generative Video Trajectory Re-Rendering using Large 4D Reconstruction Models
by: Xie, Mingyang, et al.
Published: (2026)
by: Xie, Mingyang, et al.
Published: (2026)
GaussEdit: Adaptive 3D Scene Editing with Text and Image Prompts
by: Shu, Zhenyu, et al.
Published: (2025)
by: Shu, Zhenyu, et al.
Published: (2025)
Adaptive Routing of Text-to-Image Generation Requests Between Large Cloud Model and Light-Weight Edge Model
by: Xin, Zewei, et al.
Published: (2024)
by: Xin, Zewei, et al.
Published: (2024)
Similar Items
-
WorldSpeech: A Multilingual Speech Corpus from Around the World
by: Asonitis, Antonis, et al.
Published: (2026) -
Alignment-Aware Decoding
by: Berdoz, Frédéric, et al.
Published: (2025) -
Virtual Fashion Photo-Shoots: Building a Large-Scale Garment-Lookbook Dataset
by: Hauri, Yannick, et al.
Published: (2025) -
High-Fidelity Speech Enhancement via Discrete Audio Tokens
by: Lanzendörfer, Luca A., et al.
Published: (2025) -
FLIP Reasoning Challenge
by: Plesner, Andreas, et al.
Published: (2025)