LLMs in Political Science: Heralding a New Era of Visual Analysis
Fuente:
arXiv
Guardado en:
| Autor principal: | Wang, Yu |
|---|---|
| Formato: | Preprint |
| Publicado: |
2024
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
VideoQA in the Era of LLMs: An Empirical Study
por: Xiao, Junbin, et al.
Publicado: (2024)
por: Xiao, Junbin, et al.
Publicado: (2024)
VIST-GPT: Ushering in the Era of Visual Storytelling with LLMs?
por: Gado, Mohamed, et al.
Publicado: (2025)
por: Gado, Mohamed, et al.
Publicado: (2025)
Generative Visual Communication in the Era of Vision-Language Models
por: Vinker, Yael
Publicado: (2024)
por: Vinker, Yael
Publicado: (2024)
Reasoning Path and Latent State Analysis for Multi-view Visual Spatial Reasoning: A Cognitive Science Perspective
por: Xue, Qiyao, et al.
Publicado: (2025)
por: Xue, Qiyao, et al.
Publicado: (2025)
Aligning AI with Public Values: Deliberation and Decision-Making for Governing Multimodal LLMs in Political Video Analysis
por: Sharma, Tanusree, et al.
Publicado: (2024)
por: Sharma, Tanusree, et al.
Publicado: (2024)
Visual Knowledge in the Big Model Era: Retrospect and Prospect
por: Wang, Wenguan, et al.
Publicado: (2024)
por: Wang, Wenguan, et al.
Publicado: (2024)
Multimodal LLMs Struggle with Basic Visual Network Analysis: a VNA Benchmark
por: Williams, Evan M., et al.
Publicado: (2024)
por: Williams, Evan M., et al.
Publicado: (2024)
Benchmarking Visual LLMs Resilience to Unanswerable Questions on Visually Rich Documents
por: Napolitano, Davide, et al.
Publicado: (2025)
por: Napolitano, Davide, et al.
Publicado: (2025)
Rethinking Visual Information Processing in Multimodal LLMs
por: Kim, Dongwan, et al.
Publicado: (2025)
por: Kim, Dongwan, et al.
Publicado: (2025)
UniVCD: A New Method for Unsupervised Change Detection in the Open-Vocabulary Era
por: Zhu, Ziqiang, et al.
Publicado: (2025)
por: Zhu, Ziqiang, et al.
Publicado: (2025)
AVTrustBench: Assessing and Enhancing Reliability and Robustness in Audio-Visual LLMs
por: Chowdhury, Sanjoy, et al.
Publicado: (2025)
por: Chowdhury, Sanjoy, et al.
Publicado: (2025)
Diffusion-Based Visual Art Creation: A Survey and New Perspectives
por: Wang, Bingyuan, et al.
Publicado: (2024)
por: Wang, Bingyuan, et al.
Publicado: (2024)
LatentLens: Revealing Highly Interpretable Visual Tokens in LLMs
por: Krojer, Benno, et al.
Publicado: (2026)
por: Krojer, Benno, et al.
Publicado: (2026)
Free Video-LLM: Prompt-guided Visual Perception for Efficient Training-free Video LLMs
por: Han, Kai, et al.
Publicado: (2024)
por: Han, Kai, et al.
Publicado: (2024)
Dynamic Analysis and Adaptive Discriminator for Fake News Detection
por: Su, Xinqi, et al.
Publicado: (2024)
por: Su, Xinqi, et al.
Publicado: (2024)
Euclid: Supercharging Multimodal LLMs with Synthetic High-Fidelity Visual Descriptions
por: Zhang, Jiarui, et al.
Publicado: (2024)
por: Zhang, Jiarui, et al.
Publicado: (2024)
Empowering LLMs with Pseudo-Untrimmed Videos for Audio-Visual Temporal Understanding
por: Tang, Yolo Yunlong, et al.
Publicado: (2024)
por: Tang, Yolo Yunlong, et al.
Publicado: (2024)
Chain-of-Thought Degrades Visual Spatial Reasoning Capabilities of Multimodal LLMs
por: Kancheti, Sai Srinivas, et al.
Publicado: (2026)
por: Kancheti, Sai Srinivas, et al.
Publicado: (2026)
Do MLLMs Really See It: Reinforcing Visual Attention in Multimodal LLMs
por: Ou, Siqu, et al.
Publicado: (2026)
por: Ou, Siqu, et al.
Publicado: (2026)
VISUALCENT: Visual Human Analysis using Dynamic Centroid Representation
por: Ahmad, Niaz, et al.
Publicado: (2025)
por: Ahmad, Niaz, et al.
Publicado: (2025)
Mind's Eye: A Benchmark of Visual Abstraction, Transformation and Composition for Multimodal LLMs
por: Sinha, Rohit, et al.
Publicado: (2026)
por: Sinha, Rohit, et al.
Publicado: (2026)
Generative Semantic Coding for Ultra-Low Bitrate Visual Communication and Analysis
por: Chen, Weiming, et al.
Publicado: (2025)
por: Chen, Weiming, et al.
Publicado: (2025)
Beyond Visual Memory: Mechanistic Diagnostics of Latent Visual Reasoning
por: Guo, Garvin, et al.
Publicado: (2026)
por: Guo, Garvin, et al.
Publicado: (2026)
LLMs Can Compensate for Deficiencies in Visual Representations
por: Takishita, Sho, et al.
Publicado: (2025)
por: Takishita, Sho, et al.
Publicado: (2025)
Charting New Territories: Exploring the Geographic and Geospatial Capabilities of Multimodal LLMs
por: Roberts, Jonathan, et al.
Publicado: (2023)
por: Roberts, Jonathan, et al.
Publicado: (2023)
Unlocking Attributes' Contribution to Successful Camouflage: A Combined Textual and VisualAnalysis Strategy
por: Zhang, Hong, et al.
Publicado: (2024)
por: Zhang, Hong, et al.
Publicado: (2024)
Token Activation Map to Visually Explain Multimodal LLMs
por: Li, Yi, et al.
Publicado: (2025)
por: Li, Yi, et al.
Publicado: (2025)
How Multimodal LLMs Solve Image Tasks: A Lens on Visual Grounding, Task Reasoning, and Answer Decoding
por: Yu, Zhuoran, et al.
Publicado: (2025)
por: Yu, Zhuoran, et al.
Publicado: (2025)
Sparrow: Text-Anchored Window Attention with Visual-Semantic Glimpsing for Speculative Decoding in Video LLMs
por: Zhang, Libo, et al.
Publicado: (2026)
por: Zhang, Libo, et al.
Publicado: (2026)
Affective Video Content Analysis: Decade Review and New Perspectives
por: Xue, Junxiao, et al.
Publicado: (2023)
por: Xue, Junxiao, et al.
Publicado: (2023)
Medical Reasoning in the Era of LLMs: A Systematic Review of Enhancement Techniques and Applications
por: Wang, Wenxuan, et al.
Publicado: (2025)
por: Wang, Wenxuan, et al.
Publicado: (2025)
DIVER: Dynamic Iterative Visual Evidence Reasoning for Multimodal Fake News Detection
por: Zhou, Weilin, et al.
Publicado: (2026)
por: Zhou, Weilin, et al.
Publicado: (2026)
AI-generated Image Quality Assessment in Visual Communication
por: Tian, Yu, et al.
Publicado: (2024)
por: Tian, Yu, et al.
Publicado: (2024)
Archaeoscape: Bringing Aerial Laser Scanning Archaeology to the Deep Learning Era
por: Perron, Yohann, et al.
Publicado: (2024)
por: Perron, Yohann, et al.
Publicado: (2024)
Rethinking Artistic Copyright Infringements in the Era of Text-to-Image Generative Models
por: Moayeri, Mazda, et al.
Publicado: (2024)
por: Moayeri, Mazda, et al.
Publicado: (2024)
ATSTrack: Enhancing Visual-Language Tracking by Aligning Temporal and Spatial Scales
por: Zhen, Yihao, et al.
Publicado: (2025)
por: Zhen, Yihao, et al.
Publicado: (2025)
Progressive Language-guided Visual Learning for Multi-Task Visual Grounding
por: Wang, Jingchao, et al.
Publicado: (2025)
por: Wang, Jingchao, et al.
Publicado: (2025)
Sharp Eyes and Memory for VideoLLMs: Information-Aware Visual Token Pruning for Efficient and Reliable VideoLLM Reasoning
por: Qin, Jialong, et al.
Publicado: (2025)
por: Qin, Jialong, et al.
Publicado: (2025)
Fine-R1: Make Multi-modal LLMs Excel in Fine-Grained Visual Recognition by Chain-of-Thought Reasoning
por: He, Hulingxiao, et al.
Publicado: (2026)
por: He, Hulingxiao, et al.
Publicado: (2026)
Seeing Through the Brain: New Insights from Decoding Visual Stimuli with fMRI
por: Huang, Zheng, et al.
Publicado: (2025)
por: Huang, Zheng, et al.
Publicado: (2025)
Ejemplares similares
-
VideoQA in the Era of LLMs: An Empirical Study
por: Xiao, Junbin, et al.
Publicado: (2024) -
VIST-GPT: Ushering in the Era of Visual Storytelling with LLMs?
por: Gado, Mohamed, et al.
Publicado: (2025) -
Generative Visual Communication in the Era of Vision-Language Models
por: Vinker, Yael
Publicado: (2024) -
Reasoning Path and Latent State Analysis for Multi-view Visual Spatial Reasoning: A Cognitive Science Perspective
por: Xue, Qiyao, et al.
Publicado: (2025) -
Aligning AI with Public Values: Deliberation and Decision-Making for Governing Multimodal LLMs in Political Video Analysis
por: Sharma, Tanusree, et al.
Publicado: (2024)