Visual Chronicles: Using Multimodal LLMs to Analyze Massive Collections of Images
Fuente:
arXiv
Salvato in:
| Autori principali: | Deng, Boyang, Peng, Songyou, Genova, Kyle, Wetzstein, Gordon, Snavely, Noah, Guibas, Leonidas, Funkhouser, Thomas |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2025
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
SplatTalk: 3D VQA with Gaussian Splatting
di: Thai, Anh, et al.
Pubblicazione: (2025)
di: Thai, Anh, et al.
Pubblicazione: (2025)
Streetscapes: Large-scale Consistent Street View Generation Using Autoregressive Video Diffusion
di: Deng, Boyang, et al.
Pubblicazione: (2024)
di: Deng, Boyang, et al.
Pubblicazione: (2024)
MoMaps: Semantics-Aware Scene Motion Generation with Motion Maps
di: Lei, Jiahui, et al.
Pubblicazione: (2025)
di: Lei, Jiahui, et al.
Pubblicazione: (2025)
CityRAG: Stepping Into a City via Spatially-Grounded Video Generation
di: Chou, Gene, et al.
Pubblicazione: (2026)
di: Chou, Gene, et al.
Pubblicazione: (2026)
CL-Splats: Continual Learning of Gaussian Splatting with Local Optimization
di: Ackermann, Jan, et al.
Pubblicazione: (2025)
di: Ackermann, Jan, et al.
Pubblicazione: (2025)
GeoFlow: Enforcing Implicit Geometric Consistency in Video Generation
di: Ackermann, Jan, et al.
Pubblicazione: (2026)
di: Ackermann, Jan, et al.
Pubblicazione: (2026)
Robust Symmetry Detection via Riemannian Langevin Dynamics
di: Je, Jihyeon, et al.
Pubblicazione: (2024)
di: Je, Jihyeon, et al.
Pubblicazione: (2024)
Asymmetric Flow Models
di: Chen, Hansheng, et al.
Pubblicazione: (2026)
di: Chen, Hansheng, et al.
Pubblicazione: (2026)
Policy-based Foveated Imaging and Perception
di: Xiao, Howard, et al.
Pubblicazione: (2026)
di: Xiao, Howard, et al.
Pubblicazione: (2026)
Diffusion Self-Distillation for Zero-Shot Customized Image Generation
di: Cai, Shengqu, et al.
Pubblicazione: (2024)
di: Cai, Shengqu, et al.
Pubblicazione: (2024)
Self-Calibrating Gaussian Splatting for Large Field of View Reconstruction
di: Deng, Youming, et al.
Pubblicazione: (2025)
di: Deng, Youming, et al.
Pubblicazione: (2025)
SplatPainter: Interactive Authoring of 3D Gaussians from 2D Edits via Test-Time Training
di: Zheng, Yang, et al.
Pubblicazione: (2025)
di: Zheng, Yang, et al.
Pubblicazione: (2025)
pi-Flow: Policy-Based Few-Step Generation via Imitation Distillation
di: Chen, Hansheng, et al.
Pubblicazione: (2025)
di: Chen, Hansheng, et al.
Pubblicazione: (2025)
GaussianLens: Localized High-Resolution Reconstruction via On-Demand Gaussian Densification
di: Weng, Yijia, et al.
Pubblicazione: (2025)
di: Weng, Yijia, et al.
Pubblicazione: (2025)
Generic 3D Diffusion Adapter Using Controlled Multi-View Editing
di: Chen, Hansheng, et al.
Pubblicazione: (2024)
di: Chen, Hansheng, et al.
Pubblicazione: (2024)
Dynamic Gaussian Marbles for Novel View Synthesis of Casual Monocular Videos
di: Stearns, Colton, et al.
Pubblicazione: (2024)
di: Stearns, Colton, et al.
Pubblicazione: (2024)
An AI-Powered Framework for Analyzing Collective Idea Evolution in Deliberative Assemblies
di: Poole-Dayan, Elinor, et al.
Pubblicazione: (2025)
di: Poole-Dayan, Elinor, et al.
Pubblicazione: (2025)
Taming Toxic Talk: Using chatbots to intervene with users posting toxic comments
di: Foote, Jeremy, et al.
Pubblicazione: (2026)
di: Foote, Jeremy, et al.
Pubblicazione: (2026)
Analyzing Fairness in Deepfake Detection With Massively Annotated Databases
di: Xu, Ying, et al.
Pubblicazione: (2022)
di: Xu, Ying, et al.
Pubblicazione: (2022)
ROBOPSY PL[AI]: Using Role-Play to Investigate how LLMs Present Collective Memory
di: Jahrmann, Margarete, et al.
Pubblicazione: (2025)
di: Jahrmann, Margarete, et al.
Pubblicazione: (2025)
Collaborative Video Diffusion: Consistent Multi-video Generation with Camera Control
di: Kuang, Zhengfei, et al.
Pubblicazione: (2024)
di: Kuang, Zhengfei, et al.
Pubblicazione: (2024)
Seeing Candidates at Scale: Multimodal LLMs for Visual Political Communication on Instagram
di: Achmann-Denkler, Michael, et al.
Pubblicazione: (2026)
di: Achmann-Denkler, Michael, et al.
Pubblicazione: (2026)
Using psychological theory to ground guidelines for the annotation of misogynistic language
di: Deligianni, Artemis, et al.
Pubblicazione: (2026)
di: Deligianni, Artemis, et al.
Pubblicazione: (2026)
"Amazing, They All Lean Left" -- Analyzing the Political Temperaments of Current LLMs
di: Neuman, W. Russell, et al.
Pubblicazione: (2025)
di: Neuman, W. Russell, et al.
Pubblicazione: (2025)
AIpparel: A Multimodal Foundation Model for Digital Garments
di: Nakayama, Kiyohiro, et al.
Pubblicazione: (2024)
di: Nakayama, Kiyohiro, et al.
Pubblicazione: (2024)
Listening with Language Models: Using LLMs to Collect and Interpret Classroom Feedback
di: Maram, Sai Siddartha, et al.
Pubblicazione: (2025)
di: Maram, Sai Siddartha, et al.
Pubblicazione: (2025)
Collective Recourse for Generative Urban Visualizations
di: Mushkani, Rashid
Pubblicazione: (2025)
di: Mushkani, Rashid
Pubblicazione: (2025)
GPT-4V(ision) is a Human-Aligned Evaluator for Text-to-3D Generation
di: Wu, Tong, et al.
Pubblicazione: (2024)
di: Wu, Tong, et al.
Pubblicazione: (2024)
Analyzing-Evaluating-Creating: Assessing Computational Thinking and Problem Solving in Visual Programming Domains
di: Ghosh, Ahana, et al.
Pubblicazione: (2024)
di: Ghosh, Ahana, et al.
Pubblicazione: (2024)
Crafting Synthetic Realities: Examining Visual Realism and Misinformation Potential of Photorealistic AI-Generated Images
di: Peng, Qiyao, et al.
Pubblicazione: (2024)
di: Peng, Qiyao, et al.
Pubblicazione: (2024)
GroomLight: Hybrid Inverse Rendering for Relightable Human Hair Appearance Modeling
di: Zheng, Yang, et al.
Pubblicazione: (2025)
di: Zheng, Yang, et al.
Pubblicazione: (2025)
MegaScenes: Scene-Level View Synthesis at Scale
di: Tung, Joseph, et al.
Pubblicazione: (2024)
di: Tung, Joseph, et al.
Pubblicazione: (2024)
Who and What? Using Linguistic Features and Annotator Characteristics to Analyze Annotation Variation
di: Maurer, Maximilian, et al.
Pubblicazione: (2026)
di: Maurer, Maximilian, et al.
Pubblicazione: (2026)
SuperDec: 3D Scene Decomposition with Superquadric Primitives
di: Fedele, Elisabetta, et al.
Pubblicazione: (2025)
di: Fedele, Elisabetta, et al.
Pubblicazione: (2025)
Why (not) use AI? Analyzing People's Reasoning and Conditions for AI Acceptability
di: Mun, Jimin, et al.
Pubblicazione: (2025)
di: Mun, Jimin, et al.
Pubblicazione: (2025)
Gaussian Mixture Flow Matching Models
di: Chen, Hansheng, et al.
Pubblicazione: (2025)
di: Chen, Hansheng, et al.
Pubblicazione: (2025)
Take Caution in Using LLMs as Human Surrogates: Scylla Ex Machina
di: Gao, Yuan, et al.
Pubblicazione: (2024)
di: Gao, Yuan, et al.
Pubblicazione: (2024)
ProWis: A Visual Approach for Building, Managing, and Analyzing Weather Simulation Ensembles at Runtime
di: de Souza, Carolina Veiga Ferreira, et al.
Pubblicazione: (2023)
di: de Souza, Carolina Veiga Ferreira, et al.
Pubblicazione: (2023)
Urban Mobility Assessment Using LLMs
di: Bhandari, Prabin, et al.
Pubblicazione: (2024)
di: Bhandari, Prabin, et al.
Pubblicazione: (2024)
Methods to Estimate Advanced Driver Assistance System Penetration Rates in the United States
di: Goodall, Noah
Pubblicazione: (2024)
di: Goodall, Noah
Pubblicazione: (2024)
Documenti analoghi
-
SplatTalk: 3D VQA with Gaussian Splatting
di: Thai, Anh, et al.
Pubblicazione: (2025) -
Streetscapes: Large-scale Consistent Street View Generation Using Autoregressive Video Diffusion
di: Deng, Boyang, et al.
Pubblicazione: (2024) -
MoMaps: Semantics-Aware Scene Motion Generation with Motion Maps
di: Lei, Jiahui, et al.
Pubblicazione: (2025) -
CityRAG: Stepping Into a City via Spatially-Grounded Video Generation
di: Chou, Gene, et al.
Pubblicazione: (2026) -
CL-Splats: Continual Learning of Gaussian Splatting with Local Optimization
di: Ackermann, Jan, et al.
Pubblicazione: (2025)