CoMa: Contextual Massing Generation with Vision-Language Models
Fuente:
arXiv
Saved in:
| Main Authors: | Maslov, Evgenii, Khrulkov, Valentin, Volkova, Anastasia, Gusarov, Anton, Kuznetsov, Andrey, Oseledets, Ivan |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Multi-Agent GraphRAG: A Text-to-Cypher Framework for Labeled Property Graphs
by: Gusarov, Anton, et al.
Published: (2025)
by: Gusarov, Anton, et al.
Published: (2025)
Listener-Rewarded Thinking in VLMs for Image Preferences
by: Gambashidze, Alexander, et al.
Published: (2025)
by: Gambashidze, Alexander, et al.
Published: (2025)
Test-Time Reasoning Through Visual Human Preferences with VLMs and Soft Rewards
by: Gambashidze, Alexander, et al.
Published: (2025)
by: Gambashidze, Alexander, et al.
Published: (2025)
Spread them Apart: Towards Robust Watermarking of Generated Content
by: Pautov, Mikhail, et al.
Published: (2025)
by: Pautov, Mikhail, et al.
Published: (2025)
Switti: Designing Scale-Wise Transformers for Text-to-Image Synthesis
by: Voronov, Anton, et al.
Published: (2024)
by: Voronov, Anton, et al.
Published: (2024)
General Lipschitz: Certified Robustness Against Resolvable Semantic Transformations via Transformation-Dependent Randomized Smoothing
by: Korzh, Dmitrii, et al.
Published: (2023)
by: Korzh, Dmitrii, et al.
Published: (2023)
MaxInfo: A Training-Free Key-Frame Selection Method Using Maximum Volume for Enhanced Video Understanding
by: Li, Pengyi, et al.
Published: (2025)
by: Li, Pengyi, et al.
Published: (2025)
ImageRef-VL: Enabling Contextual Image Referencing in Vision-Language Models
by: Yi, Jingwei, et al.
Published: (2025)
by: Yi, Jingwei, et al.
Published: (2025)
Random Direct Preference Optimization for Radiography Report Generation
by: Samokhin, Valentin, et al.
Published: (2025)
by: Samokhin, Valentin, et al.
Published: (2025)
CoViPAL: Layer-wise Contextualized Visual Token Pruning for Large Vision-Language Models
by: Tang, Zicong, et al.
Published: (2025)
by: Tang, Zicong, et al.
Published: (2025)
Contextual inference from single objects in Vision-Language models
by: Vilas, Martina G., et al.
Published: (2026)
by: Vilas, Martina G., et al.
Published: (2026)
Inverting Black-Box Face Recognition Systems via Zero-Order Optimization in Eigenface Space
by: Razzhigaev, Anton, et al.
Published: (2025)
by: Razzhigaev, Anton, et al.
Published: (2025)
ImprovEvolve: Ask AlphaEvolve to Improve the Input Solution and Then Improvise
by: Kravatskiy, Alexey, et al.
Published: (2026)
by: Kravatskiy, Alexey, et al.
Published: (2026)
DreamBoothDPO: Improving Personalized Generation using Direct Preference Optimization
by: Ayupov, Shamil, et al.
Published: (2025)
by: Ayupov, Shamil, et al.
Published: (2025)
PhysQuantAgent: An Inference Pipeline of Mass Estimation for Vision-Language Models
by: Yokomizo, Hisayuki, et al.
Published: (2026)
by: Yokomizo, Hisayuki, et al.
Published: (2026)
Simple Vision-Language Math Reasoning via Rendered Text
by: Skripkin, Matvey, et al.
Published: (2025)
by: Skripkin, Matvey, et al.
Published: (2025)
Speech-to-LaTeX: New Models and Datasets for Converting Spoken Equations and Sentences
by: Korzh, Dmitrii, et al.
Published: (2025)
by: Korzh, Dmitrii, et al.
Published: (2025)
VideoMaMa: Mask-Guided Video Matting via Generative Prior
by: Lim, Sangbeom, et al.
Published: (2026)
by: Lim, Sangbeom, et al.
Published: (2026)
CoFi-Dec: Hallucination-Resistant Decoding via Coarse-to-Fine Generative Feedback in Large Vision-Language Models
by: Cao, Zongsheng, et al.
Published: (2025)
by: Cao, Zongsheng, et al.
Published: (2025)
Contextual Object Detection with Multimodal Large Language Models
by: Zang, Yuhang, et al.
Published: (2023)
by: Zang, Yuhang, et al.
Published: (2023)
Generative Visual Communication in the Era of Vision-Language Models
by: Vinker, Yael
Published: (2024)
by: Vinker, Yael
Published: (2024)
Contextual Gesture: Co-Speech Gesture Video Generation through Context-aware Gesture Representation
by: Liu, Pinxin, et al.
Published: (2025)
by: Liu, Pinxin, et al.
Published: (2025)
Are Unified Vision-Language Models Necessary: Generalization Across Understanding and Generation
by: Zhang, Jihai, et al.
Published: (2025)
by: Zhang, Jihai, et al.
Published: (2025)
Distilling 3D Spatial Reasoning into a Lightweight Vision-Language Model with CoT
by: Asfour, Alaa, et al.
Published: (2026)
by: Asfour, Alaa, et al.
Published: (2026)
ClinCoT: Clinical-Aware Visual Chain-of-Thought for Medical Vision Language Models
by: Liu, Xiwei, et al.
Published: (2026)
by: Liu, Xiwei, et al.
Published: (2026)
Real-World Transferable Adversarial Attack on Face-Recognition Systems
by: Kaznacheev, Andrey, et al.
Published: (2025)
by: Kaznacheev, Andrey, et al.
Published: (2025)
Contextualized Visual Personalization in Vision-Language Models
by: Oh, Yeongtak, et al.
Published: (2026)
by: Oh, Yeongtak, et al.
Published: (2026)
Overcoming the Pitfalls of Vision-Language Model Finetuning for OOD Generalization
by: Zang, Yuhang, et al.
Published: (2024)
by: Zang, Yuhang, et al.
Published: (2024)
Dynamic Token Reduction during Generation for Vision Language Models
by: Liang, Xiaoyu, et al.
Published: (2025)
by: Liang, Xiaoyu, et al.
Published: (2025)
Florence-VL: Enhancing Vision-Language Models with Generative Vision Encoder and Depth-Breadth Fusion
by: Chen, Jiuhai, et al.
Published: (2024)
by: Chen, Jiuhai, et al.
Published: (2024)
Visual Funnel: Resolving Contextual Blindness in Multimodal Large Language Models
by: Jung, Woojun, et al.
Published: (2025)
by: Jung, Woojun, et al.
Published: (2025)
Human-Like Coarse Object Representations in Vision Models
by: Gizdov, Andrey, et al.
Published: (2026)
by: Gizdov, Andrey, et al.
Published: (2026)
Temporally-Grounded Language Generation: A Benchmark for Real-Time Vision-Language Models
by: Yu, Keunwoo Peter, et al.
Published: (2025)
by: Yu, Keunwoo Peter, et al.
Published: (2025)
3rd Workshop on Maritime Computer Vision (MaCVi) 2025: Challenge Results
by: Kiefer, Benjamin, et al.
Published: (2025)
by: Kiefer, Benjamin, et al.
Published: (2025)
CoMT: A Novel Benchmark for Chain of Multi-modal Thought on Large Vision-Language Models
by: Cheng, Zihui, et al.
Published: (2024)
by: Cheng, Zihui, et al.
Published: (2024)
How Blind and Low-Vision Individuals Prefer Large Vision-Language Model-Generated Scene Descriptions
by: An, Na Min, et al.
Published: (2025)
by: An, Na Min, et al.
Published: (2025)
IntCoOp: Interpretability-Aware Vision-Language Prompt Tuning
by: Ghosal, Soumya Suvra, et al.
Published: (2024)
by: Ghosal, Soumya Suvra, et al.
Published: (2024)
Personalized Generative Models for Contextual Debiasing
by: Liang, Xinran, et al.
Published: (2026)
by: Liang, Xinran, et al.
Published: (2026)
VLM4VLA: Revisiting Vision-Language-Models in Vision-Language-Action Models
by: Zhang, Jianke, et al.
Published: (2026)
by: Zhang, Jianke, et al.
Published: (2026)
Bias Detection and Rotation-Robustness Mitigation in Vision-Language Models and Generative Image Models
by: Mithila, Tarannum
Published: (2026)
by: Mithila, Tarannum
Published: (2026)
Similar Items
-
Multi-Agent GraphRAG: A Text-to-Cypher Framework for Labeled Property Graphs
by: Gusarov, Anton, et al.
Published: (2025) -
Listener-Rewarded Thinking in VLMs for Image Preferences
by: Gambashidze, Alexander, et al.
Published: (2025) -
Test-Time Reasoning Through Visual Human Preferences with VLMs and Soft Rewards
by: Gambashidze, Alexander, et al.
Published: (2025) -
Spread them Apart: Towards Robust Watermarking of Generated Content
by: Pautov, Mikhail, et al.
Published: (2025) -
Switti: Designing Scale-Wise Transformers for Text-to-Image Synthesis
by: Voronov, Anton, et al.
Published: (2024)