CoMa: Contextual Massing Generation with Vision-Language Models
Fuente:
arXiv
Guardado en:
| Autores principales: | Maslov, Evgenii, Khrulkov, Valentin, Volkova, Anastasia, Gusarov, Anton, Kuznetsov, Andrey, Oseledets, Ivan |
|---|---|
| Formato: | Preprint |
| Publicado: |
2026
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Multi-Agent GraphRAG: A Text-to-Cypher Framework for Labeled Property Graphs
por: Gusarov, Anton, et al.
Publicado: (2025)
por: Gusarov, Anton, et al.
Publicado: (2025)
Listener-Rewarded Thinking in VLMs for Image Preferences
por: Gambashidze, Alexander, et al.
Publicado: (2025)
por: Gambashidze, Alexander, et al.
Publicado: (2025)
Test-Time Reasoning Through Visual Human Preferences with VLMs and Soft Rewards
por: Gambashidze, Alexander, et al.
Publicado: (2025)
por: Gambashidze, Alexander, et al.
Publicado: (2025)
Spread them Apart: Towards Robust Watermarking of Generated Content
por: Pautov, Mikhail, et al.
Publicado: (2025)
por: Pautov, Mikhail, et al.
Publicado: (2025)
Switti: Designing Scale-Wise Transformers for Text-to-Image Synthesis
por: Voronov, Anton, et al.
Publicado: (2024)
por: Voronov, Anton, et al.
Publicado: (2024)
General Lipschitz: Certified Robustness Against Resolvable Semantic Transformations via Transformation-Dependent Randomized Smoothing
por: Korzh, Dmitrii, et al.
Publicado: (2023)
por: Korzh, Dmitrii, et al.
Publicado: (2023)
MaxInfo: A Training-Free Key-Frame Selection Method Using Maximum Volume for Enhanced Video Understanding
por: Li, Pengyi, et al.
Publicado: (2025)
por: Li, Pengyi, et al.
Publicado: (2025)
ImageRef-VL: Enabling Contextual Image Referencing in Vision-Language Models
por: Yi, Jingwei, et al.
Publicado: (2025)
por: Yi, Jingwei, et al.
Publicado: (2025)
Random Direct Preference Optimization for Radiography Report Generation
por: Samokhin, Valentin, et al.
Publicado: (2025)
por: Samokhin, Valentin, et al.
Publicado: (2025)
CoViPAL: Layer-wise Contextualized Visual Token Pruning for Large Vision-Language Models
por: Tang, Zicong, et al.
Publicado: (2025)
por: Tang, Zicong, et al.
Publicado: (2025)
Contextual inference from single objects in Vision-Language models
por: Vilas, Martina G., et al.
Publicado: (2026)
por: Vilas, Martina G., et al.
Publicado: (2026)
Inverting Black-Box Face Recognition Systems via Zero-Order Optimization in Eigenface Space
por: Razzhigaev, Anton, et al.
Publicado: (2025)
por: Razzhigaev, Anton, et al.
Publicado: (2025)
ImprovEvolve: Ask AlphaEvolve to Improve the Input Solution and Then Improvise
por: Kravatskiy, Alexey, et al.
Publicado: (2026)
por: Kravatskiy, Alexey, et al.
Publicado: (2026)
DreamBoothDPO: Improving Personalized Generation using Direct Preference Optimization
por: Ayupov, Shamil, et al.
Publicado: (2025)
por: Ayupov, Shamil, et al.
Publicado: (2025)
PhysQuantAgent: An Inference Pipeline of Mass Estimation for Vision-Language Models
por: Yokomizo, Hisayuki, et al.
Publicado: (2026)
por: Yokomizo, Hisayuki, et al.
Publicado: (2026)
Simple Vision-Language Math Reasoning via Rendered Text
por: Skripkin, Matvey, et al.
Publicado: (2025)
por: Skripkin, Matvey, et al.
Publicado: (2025)
Speech-to-LaTeX: New Models and Datasets for Converting Spoken Equations and Sentences
por: Korzh, Dmitrii, et al.
Publicado: (2025)
por: Korzh, Dmitrii, et al.
Publicado: (2025)
VideoMaMa: Mask-Guided Video Matting via Generative Prior
por: Lim, Sangbeom, et al.
Publicado: (2026)
por: Lim, Sangbeom, et al.
Publicado: (2026)
CoFi-Dec: Hallucination-Resistant Decoding via Coarse-to-Fine Generative Feedback in Large Vision-Language Models
por: Cao, Zongsheng, et al.
Publicado: (2025)
por: Cao, Zongsheng, et al.
Publicado: (2025)
Contextual Object Detection with Multimodal Large Language Models
por: Zang, Yuhang, et al.
Publicado: (2023)
por: Zang, Yuhang, et al.
Publicado: (2023)
Generative Visual Communication in the Era of Vision-Language Models
por: Vinker, Yael
Publicado: (2024)
por: Vinker, Yael
Publicado: (2024)
Contextual Gesture: Co-Speech Gesture Video Generation through Context-aware Gesture Representation
por: Liu, Pinxin, et al.
Publicado: (2025)
por: Liu, Pinxin, et al.
Publicado: (2025)
Are Unified Vision-Language Models Necessary: Generalization Across Understanding and Generation
por: Zhang, Jihai, et al.
Publicado: (2025)
por: Zhang, Jihai, et al.
Publicado: (2025)
Distilling 3D Spatial Reasoning into a Lightweight Vision-Language Model with CoT
por: Asfour, Alaa, et al.
Publicado: (2026)
por: Asfour, Alaa, et al.
Publicado: (2026)
ClinCoT: Clinical-Aware Visual Chain-of-Thought for Medical Vision Language Models
por: Liu, Xiwei, et al.
Publicado: (2026)
por: Liu, Xiwei, et al.
Publicado: (2026)
Real-World Transferable Adversarial Attack on Face-Recognition Systems
por: Kaznacheev, Andrey, et al.
Publicado: (2025)
por: Kaznacheev, Andrey, et al.
Publicado: (2025)
Contextualized Visual Personalization in Vision-Language Models
por: Oh, Yeongtak, et al.
Publicado: (2026)
por: Oh, Yeongtak, et al.
Publicado: (2026)
Overcoming the Pitfalls of Vision-Language Model Finetuning for OOD Generalization
por: Zang, Yuhang, et al.
Publicado: (2024)
por: Zang, Yuhang, et al.
Publicado: (2024)
Dynamic Token Reduction during Generation for Vision Language Models
por: Liang, Xiaoyu, et al.
Publicado: (2025)
por: Liang, Xiaoyu, et al.
Publicado: (2025)
Florence-VL: Enhancing Vision-Language Models with Generative Vision Encoder and Depth-Breadth Fusion
por: Chen, Jiuhai, et al.
Publicado: (2024)
por: Chen, Jiuhai, et al.
Publicado: (2024)
Visual Funnel: Resolving Contextual Blindness in Multimodal Large Language Models
por: Jung, Woojun, et al.
Publicado: (2025)
por: Jung, Woojun, et al.
Publicado: (2025)
Human-Like Coarse Object Representations in Vision Models
por: Gizdov, Andrey, et al.
Publicado: (2026)
por: Gizdov, Andrey, et al.
Publicado: (2026)
Temporally-Grounded Language Generation: A Benchmark for Real-Time Vision-Language Models
por: Yu, Keunwoo Peter, et al.
Publicado: (2025)
por: Yu, Keunwoo Peter, et al.
Publicado: (2025)
3rd Workshop on Maritime Computer Vision (MaCVi) 2025: Challenge Results
por: Kiefer, Benjamin, et al.
Publicado: (2025)
por: Kiefer, Benjamin, et al.
Publicado: (2025)
CoMT: A Novel Benchmark for Chain of Multi-modal Thought on Large Vision-Language Models
por: Cheng, Zihui, et al.
Publicado: (2024)
por: Cheng, Zihui, et al.
Publicado: (2024)
How Blind and Low-Vision Individuals Prefer Large Vision-Language Model-Generated Scene Descriptions
por: An, Na Min, et al.
Publicado: (2025)
por: An, Na Min, et al.
Publicado: (2025)
IntCoOp: Interpretability-Aware Vision-Language Prompt Tuning
por: Ghosal, Soumya Suvra, et al.
Publicado: (2024)
por: Ghosal, Soumya Suvra, et al.
Publicado: (2024)
Personalized Generative Models for Contextual Debiasing
por: Liang, Xinran, et al.
Publicado: (2026)
por: Liang, Xinran, et al.
Publicado: (2026)
VLM4VLA: Revisiting Vision-Language-Models in Vision-Language-Action Models
por: Zhang, Jianke, et al.
Publicado: (2026)
por: Zhang, Jianke, et al.
Publicado: (2026)
Bias Detection and Rotation-Robustness Mitigation in Vision-Language Models and Generative Image Models
por: Mithila, Tarannum
Publicado: (2026)
por: Mithila, Tarannum
Publicado: (2026)
Ejemplares similares
-
Multi-Agent GraphRAG: A Text-to-Cypher Framework for Labeled Property Graphs
por: Gusarov, Anton, et al.
Publicado: (2025) -
Listener-Rewarded Thinking in VLMs for Image Preferences
por: Gambashidze, Alexander, et al.
Publicado: (2025) -
Test-Time Reasoning Through Visual Human Preferences with VLMs and Soft Rewards
por: Gambashidze, Alexander, et al.
Publicado: (2025) -
Spread them Apart: Towards Robust Watermarking of Generated Content
por: Pautov, Mikhail, et al.
Publicado: (2025) -
Switti: Designing Scale-Wise Transformers for Text-to-Image Synthesis
por: Voronov, Anton, et al.
Publicado: (2024)