Speech-Synchronized Whiteboard Generation via VLM-Driven Structured Drawing Representations
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Prasad, Suraj, Mahapatra, Pinak |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
SynthPID: P&ID digitization from Topology-Preserving Synthetic Data
von: Prasad, Suraj, et al.
Veröffentlicht: (2026)
von: Prasad, Suraj, et al.
Veröffentlicht: (2026)
Replication Study: Federated Text-Driven Prompt Generation for Vision-Language Models
von: Prasad, Suraj, et al.
Veröffentlicht: (2025)
von: Prasad, Suraj, et al.
Veröffentlicht: (2025)
Suppressing VLM Hallucinations with Spectral Representation Filtering
von: Ali, Ameen, et al.
Veröffentlicht: (2025)
von: Ali, Ameen, et al.
Veröffentlicht: (2025)
SCoRe: Submodular Combinatorial Representation Learning
von: Majee, Anay, et al.
Veröffentlicht: (2023)
von: Majee, Anay, et al.
Veröffentlicht: (2023)
Draw Like an Artist: Complex Scene Generation with Diffusion Model via Composition, Painting, and Retouching
von: Liu, Minghao, et al.
Veröffentlicht: (2024)
von: Liu, Minghao, et al.
Veröffentlicht: (2024)
DocVLM: Make Your VLM an Efficient Reader
von: Nacson, Mor Shpigel, et al.
Veröffentlicht: (2024)
von: Nacson, Mor Shpigel, et al.
Veröffentlicht: (2024)
GANime: Generating Anime and Manga Character Drawings from Sketches with Deep Learning
von: Vu, Tai, et al.
Veröffentlicht: (2025)
von: Vu, Tai, et al.
Veröffentlicht: (2025)
Which Models have Perceptually-Aligned Gradients? An Explanation via Off-Manifold Robustness
von: Srinivas, Suraj, et al.
Veröffentlicht: (2023)
von: Srinivas, Suraj, et al.
Veröffentlicht: (2023)
VLM-Pruner: Buffering for Spatial Sparsity in an Efficient VLM Centrifugal Token Pruning Paradigm
von: Wu, Zhenkai, et al.
Veröffentlicht: (2025)
von: Wu, Zhenkai, et al.
Veröffentlicht: (2025)
Advancing Melanoma Diagnosis with Self-Supervised Neural Networks: Evaluating the Effectiveness of Different Techniques
von: Vusirikala, Srivishnu, et al.
Veröffentlicht: (2024)
von: Vusirikala, Srivishnu, et al.
Veröffentlicht: (2024)
Predicting Lung Disease Severity via Image-Based AQI Analysis using Deep Learning Techniques
von: Mahajan, Anvita, et al.
Veröffentlicht: (2024)
von: Mahajan, Anvita, et al.
Veröffentlicht: (2024)
DreamCube: 3D Panorama Generation via Multi-plane Synchronization
von: Huang, Yukun, et al.
Veröffentlicht: (2025)
von: Huang, Yukun, et al.
Veröffentlicht: (2025)
GRPO-RM: Fine-Tuning Representation Models via GRPO-Driven Reinforcement Learning
von: Xu, Yanchen, et al.
Veröffentlicht: (2025)
von: Xu, Yanchen, et al.
Veröffentlicht: (2025)
An Approach for Air Drawing Using Background Subtraction and Contour Extraction
von: Acharya, Ramkrishna
Veröffentlicht: (2025)
von: Acharya, Ramkrishna
Veröffentlicht: (2025)
VIBE: Can a VLM Read the Room?
von: Chakraborty, Tania, et al.
Veröffentlicht: (2025)
von: Chakraborty, Tania, et al.
Veröffentlicht: (2025)
Large VLM-based Stylized Sports Captioning
von: Dhar, Sauptik, et al.
Veröffentlicht: (2025)
von: Dhar, Sauptik, et al.
Veröffentlicht: (2025)
SuperTickets: Drawing Task-Agnostic Lottery Tickets from Supernets via Jointly Architecture Searching and Parameter Pruning
von: You, Haoran, et al.
Veröffentlicht: (2022)
von: You, Haoran, et al.
Veröffentlicht: (2022)
Structure by Architecture: Structured Representations without Regularization
von: Leeb, Felix, et al.
Veröffentlicht: (2020)
von: Leeb, Felix, et al.
Veröffentlicht: (2020)
Why Does RL Generalize Better Than SFT? A Data-Centric Perspective on VLM Post-Training
von: Lu, Aojun, et al.
Veröffentlicht: (2026)
von: Lu, Aojun, et al.
Veröffentlicht: (2026)
Designing Instance-Level Sampling Schedules via REINFORCE with James-Stein Shrinkage
von: Yu, Peiyu, et al.
Veröffentlicht: (2025)
von: Yu, Peiyu, et al.
Veröffentlicht: (2025)
SINR: Sparsity Driven Compressed Implicit Neural Representations
von: Jayasundara, Dhananjaya, et al.
Veröffentlicht: (2025)
von: Jayasundara, Dhananjaya, et al.
Veröffentlicht: (2025)
Prompt Disentanglement via Language Guidance and Representation Alignment for Domain Generalization
von: Cheng, De, et al.
Veröffentlicht: (2025)
von: Cheng, De, et al.
Veröffentlicht: (2025)
Discriminative Feature Attributions: Bridging Post Hoc Explainability and Inherent Interpretability
von: Bhalla, Usha, et al.
Veröffentlicht: (2023)
von: Bhalla, Usha, et al.
Veröffentlicht: (2023)
GLASS: Geometry-aware Local Alignment and Structure Synchronization Network for 2D-3D Registration
von: Cheng, Zhixin, et al.
Veröffentlicht: (2026)
von: Cheng, Zhixin, et al.
Veröffentlicht: (2026)
BMIP: Bi-directional Modality Interaction Prompt Learning for VLM
von: Lv, Song-Lin, et al.
Veröffentlicht: (2025)
von: Lv, Song-Lin, et al.
Veröffentlicht: (2025)
RISE: Enhancing VLM Image Annotation with Self-Supervised Reasoning
von: Hu, Suhang, et al.
Veröffentlicht: (2025)
von: Hu, Suhang, et al.
Veröffentlicht: (2025)
BendVLM: Test-Time Debiasing of Vision-Language Embeddings
von: Gerych, Walter, et al.
Veröffentlicht: (2024)
von: Gerych, Walter, et al.
Veröffentlicht: (2024)
Learning Structured Representations with Hyperbolic Embeddings
von: Sinha, Aditya, et al.
Veröffentlicht: (2024)
von: Sinha, Aditya, et al.
Veröffentlicht: (2024)
Progressive Monitoring of Generative Model Training Evolution
von: Prasad, Vidya, et al.
Veröffentlicht: (2024)
von: Prasad, Vidya, et al.
Veröffentlicht: (2024)
T3: Test-Time Model Merging in VLMs for Zero-Shot Medical Imaging Analysis
von: Imam, Raza, et al.
Veröffentlicht: (2025)
von: Imam, Raza, et al.
Veröffentlicht: (2025)
A Boundary-Metric Evaluation Protocol for Whiteboard Stroke Segmentation Under Extreme Imbalance
von: Korcynski, Nicholas
Veröffentlicht: (2026)
von: Korcynski, Nicholas
Veröffentlicht: (2026)
Rethinking Model Selection in VLM Through the Lens of Gromov-Wasserstein Distance
von: Li, Muyang, et al.
Veröffentlicht: (2026)
von: Li, Muyang, et al.
Veröffentlicht: (2026)
GraphVLM: Benchmarking Vision Language Models for Multimodal Graph Learning
von: Liu, Jiajin, et al.
Veröffentlicht: (2026)
von: Liu, Jiajin, et al.
Veröffentlicht: (2026)
HandsOnVLM: Vision-Language Models for Hand-Object Interaction Prediction
von: Bao, Chen, et al.
Veröffentlicht: (2024)
von: Bao, Chen, et al.
Veröffentlicht: (2024)
MSDS: Deep Structural Similarity with Multiscale Representation
von: Kang, Danling, et al.
Veröffentlicht: (2026)
von: Kang, Danling, et al.
Veröffentlicht: (2026)
USO: Unified Style and Subject-Driven Generation via Disentangled and Reward Learning
von: Wu, Shaojin, et al.
Veröffentlicht: (2025)
von: Wu, Shaojin, et al.
Veröffentlicht: (2025)
Riemannian Motion Generation: A Unified Framework for Human Motion Representation and Generation via Riemannian Flow Matching
von: Miao, Fangran, et al.
Veröffentlicht: (2026)
von: Miao, Fangran, et al.
Veröffentlicht: (2026)
ChemVLM: Exploring the Power of Multimodal Large Language Models in Chemistry Area
von: Li, Junxian, et al.
Veröffentlicht: (2024)
von: Li, Junxian, et al.
Veröffentlicht: (2024)
Progressive Alignment with VLM-LLM Feature to Augment Defect Classification for the ASE Dataset
von: Hsu, Chih-Chung, et al.
Veröffentlicht: (2024)
von: Hsu, Chih-Chung, et al.
Veröffentlicht: (2024)
Subject-driven Text-to-Image Generation via Preference-based Reinforcement Learning
von: Miao, Yanting, et al.
Veröffentlicht: (2024)
von: Miao, Yanting, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
SynthPID: P&ID digitization from Topology-Preserving Synthetic Data
von: Prasad, Suraj, et al.
Veröffentlicht: (2026) -
Replication Study: Federated Text-Driven Prompt Generation for Vision-Language Models
von: Prasad, Suraj, et al.
Veröffentlicht: (2025) -
Suppressing VLM Hallucinations with Spectral Representation Filtering
von: Ali, Ameen, et al.
Veröffentlicht: (2025) -
SCoRe: Submodular Combinatorial Representation Learning
von: Majee, Anay, et al.
Veröffentlicht: (2023) -
Draw Like an Artist: Complex Scene Generation with Diffusion Model via Composition, Painting, and Retouching
von: Liu, Minghao, et al.
Veröffentlicht: (2024)