Describing Images $\textit{Fast and Slow}$: Quantifying and Predicting the Variation in Human Signals during Visuo-Linguistic Processes
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Takmaz, Ece, Pezzelle, Sandro, Fernández, Raquel |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Decoding Emotions in Abstract Art: Cognitive Plausibility of CLIP in Recognizing Color-Emotion Associations
von: Widhoelzl, Hanna-Sophia, et al.
Veröffentlicht: (2024)
von: Widhoelzl, Hanna-Sophia, et al.
Veröffentlicht: (2024)
Model Merging to Maintain Language-Only Performance in Developmentally Plausible Multimodal Models
von: Takmaz, Ece, et al.
Veröffentlicht: (2025)
von: Takmaz, Ece, et al.
Veröffentlicht: (2025)
The BLA Benchmark: Investigating Basic Language Abilities of Pre-Trained Multimodal Models
von: Chen, Xinyi, et al.
Veröffentlicht: (2023)
von: Chen, Xinyi, et al.
Veröffentlicht: (2023)
Where is the multimodal goal post? On the Ability of Foundation Models to Recognize Contextually Important Moments
von: Surikuchi, Aditya K, et al.
Veröffentlicht: (2026)
von: Surikuchi, Aditya K, et al.
Veröffentlicht: (2026)
Common Objects Out of Context (COOCo): Investigating Multimodal Context and Semantic Scene Violations in Referential Communication
von: Merlo, Filippo, et al.
Veröffentlicht: (2025)
von: Merlo, Filippo, et al.
Veröffentlicht: (2025)
Not (yet) the whole story: Evaluating Visual Storytelling Requires More than Measuring Coherence, Grounding, and Repetition
von: Surikuchi, Aditya K, et al.
Veröffentlicht: (2024)
von: Surikuchi, Aditya K, et al.
Veröffentlicht: (2024)
Natural Language Generation from Visual Events: State-of-the-Art and Key Open Questions
von: Surikuchi, Aditya K, et al.
Veröffentlicht: (2025)
von: Surikuchi, Aditya K, et al.
Veröffentlicht: (2025)
Correlates of Image Memorability in Vision Encoders: Activations, Attention Entropy, Patch Uniformity and Autoencoder Losses
von: Takmaz, Ece, et al.
Veröffentlicht: (2025)
von: Takmaz, Ece, et al.
Veröffentlicht: (2025)
VL-GLUE: A Suite of Fundamental yet Challenging Visuo-Linguistic Reasoning Tasks
von: Sampat, Shailaja Keyur, et al.
Veröffentlicht: (2024)
von: Sampat, Shailaja Keyur, et al.
Veröffentlicht: (2024)
Naming, Describing, and Quantifying Visual Objects in Humans and LLMs
von: Testoni, Alberto, et al.
Veröffentlicht: (2024)
von: Testoni, Alberto, et al.
Veröffentlicht: (2024)
Emotional Theory of Mind: Bridging Fast Visual Processing with Slow Linguistic Reasoning
von: Etesam, Yasaman, et al.
Veröffentlicht: (2023)
von: Etesam, Yasaman, et al.
Veröffentlicht: (2023)
SafeLens: Deliberate and Efficient Video Guardrails with Fast-and-Slow Screening
von: Nahin, Shahriar Kabir, et al.
Veröffentlicht: (2026)
von: Nahin, Shahriar Kabir, et al.
Veröffentlicht: (2026)
Redemption Score: A Multi-Modal Evaluation Framework for Image Captioning via Distributional, Perceptual, and Linguistic Signal Triangulation
von: Dahal, Ashim, et al.
Veröffentlicht: (2025)
von: Dahal, Ashim, et al.
Veröffentlicht: (2025)
Talking Points: Describing and Localizing Pixels
von: Rusanovsky, Matan, et al.
Veröffentlicht: (2025)
von: Rusanovsky, Matan, et al.
Veröffentlicht: (2025)
Chain-of-Talkers (CoTalk): Fast Human Annotation of Dense Image Captions
von: Shen, Yijun, et al.
Veröffentlicht: (2025)
von: Shen, Yijun, et al.
Veröffentlicht: (2025)
Describing Differences in Image Sets with Natural Language
von: Dunlap, Lisa, et al.
Veröffentlicht: (2023)
von: Dunlap, Lisa, et al.
Veröffentlicht: (2023)
Fast-Slow Thinking GRPO for Large Vision-Language Model Reasoning
von: Xiao, Wenyi, et al.
Veröffentlicht: (2025)
von: Xiao, Wenyi, et al.
Veröffentlicht: (2025)
CLEVRER-Humans: Describing Physical and Causal Events the Human Way
von: Mao, Jiayuan, et al.
Veröffentlicht: (2023)
von: Mao, Jiayuan, et al.
Veröffentlicht: (2023)
SlowFast-VGen: Slow-Fast Learning for Action-Driven Long Video Generation
von: Hong, Yining, et al.
Veröffentlicht: (2024)
von: Hong, Yining, et al.
Veröffentlicht: (2024)
SlowFast-SCI: Slow-Fast Deep Unfolding Learning for Spectral Compressive Imaging
von: Zeng, Haijin, et al.
Veröffentlicht: (2025)
von: Zeng, Haijin, et al.
Veröffentlicht: (2025)
Think Twice, Click Once: Enhancing GUI Grounding via Fast and Slow Systems
von: Tang, Fei, et al.
Veröffentlicht: (2025)
von: Tang, Fei, et al.
Veröffentlicht: (2025)
Fast Prompt Alignment for Text-to-Image Generation
von: Mrini, Khalil, et al.
Veröffentlicht: (2024)
von: Mrini, Khalil, et al.
Veröffentlicht: (2024)
FlexCap: Describe Anything in Images in Controllable Detail
von: Dwibedi, Debidatta, et al.
Veröffentlicht: (2024)
von: Dwibedi, Debidatta, et al.
Veröffentlicht: (2024)
Linguistically Informed Multimodal Fusion for Vietnamese Scene-Text Image Captioning: Dataset, Graph Framework, and Phonological Attention
von: Nguyen, Nhi Ngoc-Yen, et al.
Veröffentlicht: (2026)
von: Nguyen, Nhi Ngoc-Yen, et al.
Veröffentlicht: (2026)
DescribeEarth: Describe Anything for Remote Sensing Images
von: Li, Kaiyu, et al.
Veröffentlicht: (2025)
von: Li, Kaiyu, et al.
Veröffentlicht: (2025)
VLMInferSlow: Evaluating the Efficiency Robustness of Large Vision-Language Models as a Service
von: Wang, Xiasi, et al.
Veröffentlicht: (2025)
von: Wang, Xiasi, et al.
Veröffentlicht: (2025)
Expressive and Generalizable Low-rank Adaptation for Large Models via Slow Cascaded Learning
von: Li, Siwei, et al.
Veröffentlicht: (2024)
von: Li, Siwei, et al.
Veröffentlicht: (2024)
Evaluating Linguistic Capabilities of Multimodal LLMs in the Lens of Few-Shot Learning
von: Dogan, Mustafa, et al.
Veröffentlicht: (2024)
von: Dogan, Mustafa, et al.
Veröffentlicht: (2024)
Open Vision Reasoner: Transferring Linguistic Cognitive Behavior for Visual Reasoning
von: Wei, Yana, et al.
Veröffentlicht: (2025)
von: Wei, Yana, et al.
Veröffentlicht: (2025)
MolSight: Molecular Property Prediction with Images
von: Baranwal, Aaditya, et al.
Veröffentlicht: (2026)
von: Baranwal, Aaditya, et al.
Veröffentlicht: (2026)
Visual Merit or Linguistic Crutch? A Close Look at DeepSeek-OCR
von: Liang, Yunhao, et al.
Veröffentlicht: (2026)
von: Liang, Yunhao, et al.
Veröffentlicht: (2026)
NAVCON: A Cognitively Inspired and Linguistically Grounded Corpus for Vision and Language Navigation
von: Wanchoo, Karan, et al.
Veröffentlicht: (2024)
von: Wanchoo, Karan, et al.
Veröffentlicht: (2024)
RadDiff: Describing Differences in Radiology Image Sets with Natural Language
von: Shen, Xiaoxian, et al.
Veröffentlicht: (2026)
von: Shen, Xiaoxian, et al.
Veröffentlicht: (2026)
Vision-Language Models Align with Human Neural Representations in Concept Processing
von: Bavaresco, Anna, et al.
Veröffentlicht: (2024)
von: Bavaresco, Anna, et al.
Veröffentlicht: (2024)
Linguistic Binding in Diffusion Models: Enhancing Attribute Correspondence through Attention Map Alignment
von: Rassin, Royi, et al.
Veröffentlicht: (2023)
von: Rassin, Royi, et al.
Veröffentlicht: (2023)
Improving Image Captioning by Mimicking Human Reformulation Feedback at Inference-time
von: Berger, Uri, et al.
Veröffentlicht: (2025)
von: Berger, Uri, et al.
Veröffentlicht: (2025)
WiFi-GEN: High-Resolution Indoor Imaging from WiFi Signals Using Generative AI
von: Shi, Jianyang, et al.
Veröffentlicht: (2024)
von: Shi, Jianyang, et al.
Veröffentlicht: (2024)
Describe Anything in Medical Images
von: Xiao, Xi, et al.
Veröffentlicht: (2025)
von: Xiao, Xi, et al.
Veröffentlicht: (2025)
FastPerson: Enhancing Video Learning through Effective Video Summarization that Preserves Linguistic and Visual Contexts
von: Kawamura, Kazuki, et al.
Veröffentlicht: (2024)
von: Kawamura, Kazuki, et al.
Veröffentlicht: (2024)
Arbitration Failure, Not Perceptual Blindness: How Vision-Language Models Resolve Visual-Linguistic Conflicts
von: Nooralahzadeh, Farhad, et al.
Veröffentlicht: (2026)
von: Nooralahzadeh, Farhad, et al.
Veröffentlicht: (2026)
Ähnliche Einträge
-
Decoding Emotions in Abstract Art: Cognitive Plausibility of CLIP in Recognizing Color-Emotion Associations
von: Widhoelzl, Hanna-Sophia, et al.
Veröffentlicht: (2024) -
Model Merging to Maintain Language-Only Performance in Developmentally Plausible Multimodal Models
von: Takmaz, Ece, et al.
Veröffentlicht: (2025) -
The BLA Benchmark: Investigating Basic Language Abilities of Pre-Trained Multimodal Models
von: Chen, Xinyi, et al.
Veröffentlicht: (2023) -
Where is the multimodal goal post? On the Ability of Foundation Models to Recognize Contextually Important Moments
von: Surikuchi, Aditya K, et al.
Veröffentlicht: (2026) -
Common Objects Out of Context (COOCo): Investigating Multimodal Context and Semantic Scene Violations in Referential Communication
von: Merlo, Filippo, et al.
Veröffentlicht: (2025)