Semantic Document Derendering: SVG Reconstruction via Vision-Language Modeling
Fuente:
arXiv
Saved in:
| Main Authors: | Hazimeh, Adam, Wang, Ke, Collier, Mark, Baechler, Gilles, Kokiopoulou, Efi, Frossard, Pascal |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Representing Online Handwriting for Recognition in Large Vision-Language Models
by: Fadeeva, Anastasiia, et al.
Published: (2024)
by: Fadeeva, Anastasiia, et al.
Published: (2024)
Pi-DUAL: Using Privileged Information to Distinguish Clean from Noisy Labels
by: Wang, Ke, et al.
Published: (2023)
by: Wang, Ke, et al.
Published: (2023)
VFIG: Vectorizing Complex Figures in SVG with Vision-Language Models
by: He, Qijia, et al.
Published: (2026)
by: He, Qijia, et al.
Published: (2026)
Sketch-to-Layout: Sketch-Guided Multimodal Layout Generation
by: Brioschi, Riccardo, et al.
Published: (2025)
by: Brioschi, Riccardo, et al.
Published: (2025)
Single-Input Multi-Output Model Merging: Leveraging Foundation Models for Dense Multi-Task Learning
by: Giraldo, Juan Garcia, et al.
Published: (2025)
by: Giraldo, Juan Garcia, et al.
Published: (2025)
ScreenAI: A Vision-Language Model for UI and Infographics Understanding
by: Baechler, Gilles, et al.
Published: (2024)
by: Baechler, Gilles, et al.
Published: (2024)
SVGDreamer: Text Guided SVG Generation with Diffusion Model
by: Xing, Ximing, et al.
Published: (2023)
by: Xing, Ximing, et al.
Published: (2023)
Sequential Representation Learning via Static-Dynamic Conditional Disentanglement
by: Simon, Mathieu Cyrille, et al.
Published: (2024)
by: Simon, Mathieu Cyrille, et al.
Published: (2024)
UniSVG: A Unified Dataset for Vector Graphic Understanding and Generation with Multimodal Large Language Models
by: Li, Jinke, et al.
Published: (2025)
by: Li, Jinke, et al.
Published: (2025)
SVGenius: Benchmarking LLMs in SVG Understanding, Editing and Generation
by: Chen, Siqi, et al.
Published: (2025)
by: Chen, Siqi, et al.
Published: (2025)
SVGDreamer++: Advancing Editability and Diversity in Text-Guided SVG Generation
by: Xing, Ximing, et al.
Published: (2024)
by: Xing, Ximing, et al.
Published: (2024)
Object-Level Verbalized Confidence Calibration in Vision-Language Models via Semantic Perturbation
by: Zhao, Yunpu, et al.
Published: (2025)
by: Zhao, Yunpu, et al.
Published: (2025)
SVGauge: Towards Human-Aligned Evaluation for SVG Generation
by: Zini, Leonardo, et al.
Published: (2025)
by: Zini, Leonardo, et al.
Published: (2025)
From Tokens to Numbers: Continuous Number Modeling for SVG Generation
by: Ogezi, Michael, et al.
Published: (2026)
by: Ogezi, Michael, et al.
Published: (2026)
DistractMIA: Black-Box Membership Inference on Vision-Language Models via Semantic Distraction
by: Tang, Hongyi, et al.
Published: (2026)
by: Tang, Hongyi, et al.
Published: (2026)
MURE: Hierarchical Multi-Resolution Encoding via Vision-Language Models for Visual Document Retrieval
by: Zhu, Fengbin, et al.
Published: (2026)
by: Zhu, Fengbin, et al.
Published: (2026)
SemVideo: Reconstructs What You Watch from Brain Activity via Hierarchical Semantic Guidance
by: Yang, Minghan, et al.
Published: (2026)
by: Yang, Minghan, et al.
Published: (2026)
SEAM: Semantically Equivalent Across Modalities Benchmark for Vision-Language Models
by: Tang, Zhenwei, et al.
Published: (2025)
by: Tang, Zhenwei, et al.
Published: (2025)
On the Utility of Foundation Models for Fast MRI: Vision-Language-Guided Image Reconstruction
by: Feng, Ruimin, et al.
Published: (2025)
by: Feng, Ruimin, et al.
Published: (2025)
SARVLM: A Vision Language Foundation Model for Semantic Understanding in SAR Imagery
by: Ma, Qiwei, et al.
Published: (2025)
by: Ma, Qiwei, et al.
Published: (2025)
SVGBuilder: Component-Based Colored SVG Generation with Text-Guided Autoregressive Transformers
by: Chen, Zehao, et al.
Published: (2024)
by: Chen, Zehao, et al.
Published: (2024)
VectorGym: A Multitask Benchmark for SVG Code Generation, Sketching, and Editing
by: Rodriguez, Juan, et al.
Published: (2026)
by: Rodriguez, Juan, et al.
Published: (2026)
Remodeling Semantic Relationships in Vision-Language Fine-Tuning
by: Wu, Xiangyang, et al.
Published: (2025)
by: Wu, Xiangyang, et al.
Published: (2025)
Document Haystacks: Vision-Language Reasoning Over Piles of 1000+ Documents
by: Chen, Jun, et al.
Published: (2024)
by: Chen, Jun, et al.
Published: (2024)
Vision-Language Model Purified Semi-Supervised Semantic Segmentation for Remote Sensing Images
by: Wang, Shanwen, et al.
Published: (2026)
by: Wang, Shanwen, et al.
Published: (2026)
Conjugated Semantic Pool Improves OOD Detection with Pre-trained Vision-Language Models
by: Chen, Mengyuan, et al.
Published: (2024)
by: Chen, Mengyuan, et al.
Published: (2024)
Emergent Open-Vocabulary Semantic Segmentation from Off-the-shelf Vision-Language Models
by: Luo, Jiayun, et al.
Published: (2023)
by: Luo, Jiayun, et al.
Published: (2023)
Evading Visual Aphasia: Contrastive Adaptive Semantic Token Pruning for Vision-Language Models
by: Ma, Jie, et al.
Published: (2026)
by: Ma, Jie, et al.
Published: (2026)
A Study on Unsupervised Domain Adaptation for Semantic Segmentation in the Era of Vision-Language Models
by: Schwonberg, Manuel, et al.
Published: (2024)
by: Schwonberg, Manuel, et al.
Published: (2024)
Single-Sample Black-Box Membership Inference Attack against Vision-Language Models via Cross-modal Semantic Alignment
by: Li, Jiaqing, et al.
Published: (2026)
by: Li, Jiaqing, et al.
Published: (2026)
Intersectional Fairness in Vision-Language Models for Medical Image Disease Classification
by: Zhang, Yupeng, et al.
Published: (2025)
by: Zhang, Yupeng, et al.
Published: (2025)
PROSPECT: Unified Streaming Vision-Language Navigation via Semantic--Spatial Fusion and Latent Predictive Representation
by: Fan, Zehua, et al.
Published: (2026)
by: Fan, Zehua, et al.
Published: (2026)
Jailbreaks on Vision Language Model via Multimodal Reasoning
by: Noheria, Aarush, et al.
Published: (2026)
by: Noheria, Aarush, et al.
Published: (2026)
A Visual Semantic Adaptive Watermark grounded by Prefix-Tuning for Large Vision-Language Model
by: Zheng, Qi, et al.
Published: (2026)
by: Zheng, Qi, et al.
Published: (2026)
LLaVA-SG: Leveraging Scene Graphs as Visual Semantic Expression in Vision-Language Models
by: Wang, Jingyi, et al.
Published: (2024)
by: Wang, Jingyi, et al.
Published: (2024)
Uncertainty-guided Compositional Alignment with Part-to-Whole Semantic Representativeness in Hyperbolic Vision-Language Models
by: Kim, Hayeon, et al.
Published: (2026)
by: Kim, Hayeon, et al.
Published: (2026)
OmniPT: Unleashing the Potential of Large Vision Language Models for Pedestrian Tracking and Understanding
by: Fu, Teng, et al.
Published: (2025)
by: Fu, Teng, et al.
Published: (2025)
Offline Semantic Guidance for Efficient Vision-Language-Action Policy Distillation
by: Shi, Jin, et al.
Published: (2026)
by: Shi, Jin, et al.
Published: (2026)
Compound Expression Recognition via Large Vision-Language Models
by: Yu, Jun, et al.
Published: (2025)
by: Yu, Jun, et al.
Published: (2025)
Safe Vision-Language Models via Unsafe Weights Manipulation
by: D'Incà, Moreno, et al.
Published: (2025)
by: D'Incà, Moreno, et al.
Published: (2025)
Similar Items
-
Representing Online Handwriting for Recognition in Large Vision-Language Models
by: Fadeeva, Anastasiia, et al.
Published: (2024) -
Pi-DUAL: Using Privileged Information to Distinguish Clean from Noisy Labels
by: Wang, Ke, et al.
Published: (2023) -
VFIG: Vectorizing Complex Figures in SVG with Vision-Language Models
by: He, Qijia, et al.
Published: (2026) -
Sketch-to-Layout: Sketch-Guided Multimodal Layout Generation
by: Brioschi, Riccardo, et al.
Published: (2025) -
Single-Input Multi-Output Model Merging: Leveraging Foundation Models for Dense Multi-Task Learning
by: Giraldo, Juan Garcia, et al.
Published: (2025)