Talking Points: Describing and Localizing Pixels
Fuente:
arXiv
Saved in:
| Main Authors: | Rusanovsky, Matan, Malnick, Shimon, Avidan, Shai |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Memories of Forgotten Concepts
by: Rusanovsky, Matan, et al.
Published: (2024)
by: Rusanovsky, Matan, et al.
Published: (2024)
CapeX: Category-Agnostic Pose Estimation from Textual Point Explanation
by: Rusanovsky, Matan, et al.
Published: (2024)
by: Rusanovsky, Matan, et al.
Published: (2024)
A Graph-Based Approach for Category-Agnostic Pose Estimation
by: Hirschorn, Or, et al.
Published: (2023)
by: Hirschorn, Or, et al.
Published: (2023)
Edge Weight Prediction For Category-Agnostic Pose Estimation
by: Hirschorn, Or, et al.
Published: (2024)
by: Hirschorn, Or, et al.
Published: (2024)
Optimize the Unseen -- Fast NeRF Cleanup with Free Space Prior
by: Segre, Leo, et al.
Published: (2024)
by: Segre, Leo, et al.
Published: (2024)
VF-NeRF: Viewshed Fields for Rigid NeRF Registration
by: Segre, Leo, et al.
Published: (2024)
by: Segre, Leo, et al.
Published: (2024)
Multi-View Foundation Models
by: Segre, Leo, et al.
Published: (2025)
by: Segre, Leo, et al.
Published: (2025)
Hearing the Room Through the Shape of the Drum: Modal-Guided Sound Recovery from Multi-Point Surface Vibrations
by: Bagon, Shai, et al.
Published: (2026)
by: Bagon, Shai, et al.
Published: (2026)
Frequency-Aware Gaussian Splatting Decomposition
by: Lavi, Yishai, et al.
Published: (2025)
by: Lavi, Yishai, et al.
Published: (2025)
Securing Neural Networks with Knapsack Optimization
by: Gorski, Yakir, et al.
Published: (2023)
by: Gorski, Yakir, et al.
Published: (2023)
Coordinate Descent for Network Linearization
by: Rakhlin, Vlad, et al.
Published: (2025)
by: Rakhlin, Vlad, et al.
Published: (2025)
PointAD: Comprehending 3D Anomalies from Points and Pixels for Zero-shot 3D Anomaly Detection
by: Zhou, Qihang, et al.
Published: (2024)
by: Zhou, Qihang, et al.
Published: (2024)
PixelWorld: How Far Are We from Perceiving Everything as Pixels?
by: Lyu, Zhiheng, et al.
Published: (2025)
by: Lyu, Zhiheng, et al.
Published: (2025)
Pixel Sentence Representation Learning
by: Xiao, Chenghao, et al.
Published: (2024)
by: Xiao, Chenghao, et al.
Published: (2024)
Autoregressive Pre-Training on Pixels and Texts
by: Chai, Yekun, et al.
Published: (2024)
by: Chai, Yekun, et al.
Published: (2024)
Scene Grounding In the Wild
by: Cohen, Tamir, et al.
Published: (2026)
by: Cohen, Tamir, et al.
Published: (2026)
Learning to See Inside Opaque Liquid Containers using Speckle Vibrometry
by: Kichler, Matan, et al.
Published: (2025)
by: Kichler, Matan, et al.
Published: (2025)
Pixel-Level Reasoning Segmentation via Multi-turn Conversations
by: Cai, Dexian, et al.
Published: (2025)
by: Cai, Dexian, et al.
Published: (2025)
From Text to Pixel: Advancing Long-Context Understanding in MLLMs
by: Lu, Yujie, et al.
Published: (2024)
by: Lu, Yujie, et al.
Published: (2024)
From Reasoning to Pixels: Benchmarking the Alignment Gap in Unified Multimodal Models
by: Yang, Cheng, et al.
Published: (2026)
by: Yang, Cheng, et al.
Published: (2026)
Chain-of-Talkers (CoTalk): Fast Human Annotation of Dense Image Captions
by: Shen, Yijun, et al.
Published: (2025)
by: Shen, Yijun, et al.
Published: (2025)
From Pixels to Tokens: Revisiting Object Hallucinations in Large Vision-Language Models
by: Shang, Yuying, et al.
Published: (2024)
by: Shang, Yuying, et al.
Published: (2024)
Describing Differences in Image Sets with Natural Language
by: Dunlap, Lisa, et al.
Published: (2023)
by: Dunlap, Lisa, et al.
Published: (2023)
Pixel Reasoner: Incentivizing Pixel-Space Reasoning with Curiosity-Driven Reinforcement Learning
by: Wang, Haozhe, et al.
Published: (2025)
by: Wang, Haozhe, et al.
Published: (2025)
Attention, Please! PixelSHAP Reveals What Vision-Language Models Actually Focus On
by: Goldshmidt, Roni
Published: (2025)
by: Goldshmidt, Roni
Published: (2025)
TextureSAM: Towards a Texture Aware Foundation Model for Segmentation
by: Cohen, Inbal, et al.
Published: (2025)
by: Cohen, Inbal, et al.
Published: (2025)
PixelPrune: Pixel-Level Adaptive Visual Token Reduction via Predictive Coding
by: Wang, Nan, et al.
Published: (2026)
by: Wang, Nan, et al.
Published: (2026)
Reading, Not Thinking: Understanding and Bridging the Modality Gap When Text Becomes Pixels in Multimodal LLMs
by: Sun, Kaiser, et al.
Published: (2026)
by: Sun, Kaiser, et al.
Published: (2026)
PinPoint: Prompting with Informative Interior Points
by: Sadeghi, Pouya, et al.
Published: (2026)
by: Sadeghi, Pouya, et al.
Published: (2026)
Catch Me If You Can Describe Me: Open-Vocabulary Camouflaged Instance Segmentation with Diffusion
by: Vu, Tuan-Anh, et al.
Published: (2023)
by: Vu, Tuan-Anh, et al.
Published: (2023)
Pixels, Patterns, but No Poetry: To See The World like Humans
by: Gao, Hongcheng, et al.
Published: (2025)
by: Gao, Hongcheng, et al.
Published: (2025)
Describing Images $\textit{Fast and Slow}$: Quantifying and Predicting the Variation in Human Signals during Visuo-Linguistic Processes
by: Takmaz, Ece, et al.
Published: (2024)
by: Takmaz, Ece, et al.
Published: (2024)
FlexCap: Describe Anything in Images in Controllable Detail
by: Dwibedi, Debidatta, et al.
Published: (2024)
by: Dwibedi, Debidatta, et al.
Published: (2024)
InternLM-XComposer2-4KHD: A Pioneering Large Vision-Language Model Handling Resolutions from 336 Pixels to 4K HD
by: Dong, Xiaoyi, et al.
Published: (2024)
by: Dong, Xiaoyi, et al.
Published: (2024)
CLEVRER-Humans: Describing Physical and Causal Events the Human Way
by: Mao, Jiayuan, et al.
Published: (2023)
by: Mao, Jiayuan, et al.
Published: (2023)
Humor in Pixels: Benchmarking Large Multimodal Models Understanding of Online Comics
by: Ryan, Yuriel, et al.
Published: (2025)
by: Ryan, Yuriel, et al.
Published: (2025)
From Pixels to Posts: Retrieval-Augmented Fashion Captioning and Hashtag Generation
by: Gondal, Moazzam Umer, et al.
Published: (2025)
by: Gondal, Moazzam Umer, et al.
Published: (2025)
From Pampas to Pixels: Fine-Tuning Diffusion Models for Gaúcho Heritage
by: Amadeus, Marcellus, et al.
Published: (2024)
by: Amadeus, Marcellus, et al.
Published: (2024)
From Pixels to Prose: A Large Dataset of Dense Image Captions
by: Singla, Vasu, et al.
Published: (2024)
by: Singla, Vasu, et al.
Published: (2024)
From Pixels to Tokens: Byte-Pair Encoding on Quantized Visual Modalities
by: Zhang, Wanpeng, et al.
Published: (2024)
by: Zhang, Wanpeng, et al.
Published: (2024)
Similar Items
-
Memories of Forgotten Concepts
by: Rusanovsky, Matan, et al.
Published: (2024) -
CapeX: Category-Agnostic Pose Estimation from Textual Point Explanation
by: Rusanovsky, Matan, et al.
Published: (2024) -
A Graph-Based Approach for Category-Agnostic Pose Estimation
by: Hirschorn, Or, et al.
Published: (2023) -
Edge Weight Prediction For Category-Agnostic Pose Estimation
by: Hirschorn, Or, et al.
Published: (2024) -
Optimize the Unseen -- Fast NeRF Cleanup with Free Space Prior
by: Segre, Leo, et al.
Published: (2024)