MarkushGrapher: Joint Visual and Textual Recognition of Markush Structures
Fuente:
arXiv
Guardado en:
| Autores principales: | Morin, Lucas, Weber, Valéry, Nassar, Ahmed, Meijer, Gerhard Ingmar, Van Gool, Luc, Li, Yawei, Staar, Peter |
|---|---|
| Formato: | Preprint |
| Publicado: |
2025
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
MarkushGrapher-2: End-to-end Multimodal Recognition of Chemical Structures
por: Strohmeyer, Tim, et al.
Publicado: (2026)
por: Strohmeyer, Tim, et al.
Publicado: (2026)
SubGrapher: Visual Fingerprinting of Chemical Structures
por: Morin, Lucas, et al.
Publicado: (2025)
por: Morin, Lucas, et al.
Publicado: (2025)
MolGrapher: Graph-based Visual Recognition of Chemical Structures
por: Morin, Lucas, et al.
Publicado: (2023)
por: Morin, Lucas, et al.
Publicado: (2023)
Маркуш Олександр Іванович [Markush Oleksandr Ivanovych]
por: В. В. Ґабор
Publicado: (2018)
por: В. В. Ґабор
Publicado: (2018)
ScreenParse: Moving Beyond Sparse Grounding with Complete Screen Parsing Supervision
por: Gurbuz, A. Said, et al.
Publicado: (2026)
por: Gurbuz, A. Said, et al.
Publicado: (2026)
Shapley Pruning for Neural Network Compression
por: Adamczewski, Kamil, et al.
Publicado: (2024)
por: Adamczewski, Kamil, et al.
Publicado: (2024)
Advanced Layout Analysis Models for Docling
por: Livathinos, Nikolaos, et al.
Publicado: (2025)
por: Livathinos, Nikolaos, et al.
Publicado: (2025)
Docling: An Efficient Open-Source Toolkit for AI-driven Document Conversion
por: Livathinos, Nikolaos, et al.
Publicado: (2025)
por: Livathinos, Nikolaos, et al.
Publicado: (2025)
Taming CLIP for Fine-grained and Structured Visual Understanding of Museum Exhibits
por: Balauca, Ada-Astrid, et al.
Publicado: (2024)
por: Balauca, Ada-Astrid, et al.
Publicado: (2024)
Fine-Grained Spatial and Verbal Losses for 3D Visual Grounding
por: Dey, Sombit, et al.
Publicado: (2024)
por: Dey, Sombit, et al.
Publicado: (2024)
Docling Technical Report
por: Auer, Christoph, et al.
Publicado: (2024)
por: Auer, Christoph, et al.
Publicado: (2024)
Four Ways to Improve Verbo-visual Fusion for Dense 3D Visual Grounding
por: Unal, Ozan, et al.
Publicado: (2023)
por: Unal, Ozan, et al.
Publicado: (2023)
LocalViT: Analyzing Locality in Vision Transformers
por: Li, Yawei, et al.
Publicado: (2021)
por: Li, Yawei, et al.
Publicado: (2021)
Visual and Textual Prompts in VLLMs for Enhancing Emotion Recognition
por: Wang, Zhifeng, et al.
Publicado: (2025)
por: Wang, Zhifeng, et al.
Publicado: (2025)
Test-time Training for Hyperspectral Image Super-resolution
por: Li, Ke, et al.
Publicado: (2024)
por: Li, Ke, et al.
Publicado: (2024)
Optimizing against Infeasible Inclusions from Data for Semantic Segmentation through Morphology
por: Basu, Shamik, et al.
Publicado: (2024)
por: Basu, Shamik, et al.
Publicado: (2024)
Bayesian Self-Training for Semi-Supervised 3D Segmentation
por: Unal, Ozan, et al.
Publicado: (2024)
por: Unal, Ozan, et al.
Publicado: (2024)
TrafficBots V1.5: Traffic Simulation via Conditional VAEs and Transformers with Relative Pose Encoding
por: Zhang, Zhejun, et al.
Publicado: (2024)
por: Zhang, Zhejun, et al.
Publicado: (2024)
Investigating the Effectiveness of Cross-Attention to Unlock Zero-Shot Editing of Text-to-Video Diffusion Models
por: Motamed, Saman, et al.
Publicado: (2024)
por: Motamed, Saman, et al.
Publicado: (2024)
EvenNICER-SLAM: Event-based Neural Implicit Encoding SLAM
por: Chen, Shi, et al.
Publicado: (2024)
por: Chen, Shi, et al.
Publicado: (2024)
SmolDocling: An ultra-compact vision-language model for end-to-end multi-modal document conversion
por: Nassar, Ahmed, et al.
Publicado: (2025)
por: Nassar, Ahmed, et al.
Publicado: (2025)
ReVLA: Reverting Visual Domain Limitation of Robotic Foundation Models
por: Dey, Sombit, et al.
Publicado: (2024)
por: Dey, Sombit, et al.
Publicado: (2024)
Enhanced Multi-Scale Cross-Attention for Person Image Generation
por: Tang, Hao, et al.
Publicado: (2025)
por: Tang, Hao, et al.
Publicado: (2025)
Towards Online Real-Time Memory-based Video Inpainting Transformers
por: Thiry, Guillaume, et al.
Publicado: (2024)
por: Thiry, Guillaume, et al.
Publicado: (2024)
Condition-Invariant Semantic Segmentation
por: Sakaridis, Christos, et al.
Publicado: (2023)
por: Sakaridis, Christos, et al.
Publicado: (2023)
Graph Transformer GANs with Graph Masked Modeling for Architectural Layout Generation
por: Tang, Hao, et al.
Publicado: (2024)
por: Tang, Hao, et al.
Publicado: (2024)
MatIR: A Hybrid Mamba-Transformer Image Restoration Model
por: Wen, Juan, et al.
Publicado: (2025)
por: Wen, Juan, et al.
Publicado: (2025)
CAFuser: Condition-Aware Multimodal Fusion for Robust Semantic Perception of Driving Scenes
por: Broedermann, Tim, et al.
Publicado: (2024)
por: Broedermann, Tim, et al.
Publicado: (2024)
Sun Off, Lights On: Photorealistic Monocular Nighttime Simulation for Robust Semantic Perception
por: Tzevelekakis, Konstantinos, et al.
Publicado: (2024)
por: Tzevelekakis, Konstantinos, et al.
Publicado: (2024)
A Simple and Generalist Approach for Panoptic Segmentation
por: Prisadnikov, Nedyalko, et al.
Publicado: (2024)
por: Prisadnikov, Nedyalko, et al.
Publicado: (2024)
ESG Accountability Made Easy: DocQA at Your Service
por: Mishra, Lokesh, et al.
Publicado: (2023)
por: Mishra, Lokesh, et al.
Publicado: (2023)
Empowering Image Recovery_ A Multi-Attention Approach
por: Wen, Juan, et al.
Publicado: (2024)
por: Wen, Juan, et al.
Publicado: (2024)
Implicit-Zoo: A Large-Scale Dataset of Neural Implicit Functions for 2D Images and 3D Scenes
por: Ma, Qi, et al.
Publicado: (2024)
por: Ma, Qi, et al.
Publicado: (2024)
Continuous Pose for Monocular Cameras in Neural Implicit Representation
por: Ma, Qi, et al.
Publicado: (2023)
por: Ma, Qi, et al.
Publicado: (2023)
From Synchrony to Sequence: Exo-to-Ego Generation via Interpolation
por: Mahdi, Mohammad, et al.
Publicado: (2026)
por: Mahdi, Mohammad, et al.
Publicado: (2026)
Vision encoders should be image size agnostic and task driven
por: Prisadnikov, Nedyalko, et al.
Publicado: (2025)
por: Prisadnikov, Nedyalko, et al.
Publicado: (2025)
Self-supervised pretraining for an iterative image size agnostic vision transformer
por: Prisadnikov, Nedyalko, et al.
Publicado: (2026)
por: Prisadnikov, Nedyalko, et al.
Publicado: (2026)
HandDiff: 3D Hand Pose Estimation with Diffusion on Image-Point Cloud
por: Cheng, Wencan, et al.
Publicado: (2024)
por: Cheng, Wencan, et al.
Publicado: (2024)
Bringing Masked Autoencoders Explicit Contrastive Properties for Point Cloud Self-Supervised Learning
por: Ren, Bin, et al.
Publicado: (2024)
por: Ren, Bin, et al.
Publicado: (2024)
Sharing Key Semantics in Transformer Makes Efficient Image Restoration
por: Ren, Bin, et al.
Publicado: (2024)
por: Ren, Bin, et al.
Publicado: (2024)
Ejemplares similares
-
MarkushGrapher-2: End-to-end Multimodal Recognition of Chemical Structures
por: Strohmeyer, Tim, et al.
Publicado: (2026) -
SubGrapher: Visual Fingerprinting of Chemical Structures
por: Morin, Lucas, et al.
Publicado: (2025) -
MolGrapher: Graph-based Visual Recognition of Chemical Structures
por: Morin, Lucas, et al.
Publicado: (2023) -
Маркуш Олександр Іванович [Markush Oleksandr Ivanovych]
por: В. В. Ґабор
Publicado: (2018) -
ScreenParse: Moving Beyond Sparse Grounding with Complete Screen Parsing Supervision
por: Gurbuz, A. Said, et al.
Publicado: (2026)