Saved in:
| Main Authors: | Pancerz, Krzysztof, Kulicki, Piotr, Kalisz, Michał, Burda, Andrzej, Stanisławski, Maciej, Sarzyński, Jaromir |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | https://arxiv.org/abs/2504.13150 |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Visible Yet Unreadable: A Systematic Blind Spot of Vision Language Models Across Writing Systems
by: Zhang, Jie, et al.
Published: (2025)
by: Zhang, Jie, et al.
Published: (2025)
Beyond Recognition: Evaluating Visual Perspective Taking in Vision Language Models
by: Góral, Gracjan, et al.
Published: (2025)
by: Góral, Gracjan, et al.
Published: (2025)
How Culturally Aware are Vision-Language Models?
by: Burda-Lassen, Olena, et al.
Published: (2024)
by: Burda-Lassen, Olena, et al.
Published: (2024)
Iconographic Classification and Content-Based Recommendation for Digitized Artworks
by: Kutt, Krzysztof, et al.
Published: (2026)
by: Kutt, Krzysztof, et al.
Published: (2026)
A Python toolkit for dealing with Petri nets over ontological graphs
by: Pancerz, Krzysztof
Published: (2025)
by: Pancerz, Krzysztof
Published: (2025)
Modeling Retinal Ganglion Cells with Neural Differential Equations
by: Dobek, Kacper, et al.
Published: (2025)
by: Dobek, Kacper, et al.
Published: (2025)
Emotion Recognition with Facial Attention and Objective Activation Functions
by: Miskow, Andrzej, et al.
Published: (2024)
by: Miskow, Andrzej, et al.
Published: (2024)
Seeing the Signs: A Survey of Edge-Deployable OCR Models for Billboard Visibility Analysis
by: Szankin, Maciej, et al.
Published: (2025)
by: Szankin, Maciej, et al.
Published: (2025)
Red-Teaming Segment Anything Model
by: Jankowski, Krzysztof, et al.
Published: (2024)
by: Jankowski, Krzysztof, et al.
Published: (2024)
DinoTwins: Combining DINO and Barlow Twins for Robust, Label-Efficient Vision Transformers
by: Podsiadly, Michael, et al.
Published: (2025)
by: Podsiadly, Michael, et al.
Published: (2025)
Exploring the Use of Contrastive Language-Image Pre-Training for Human Posture Classification: Insights from Yoga Pose Analysis
by: Dobrzycki, Andrzej D., et al.
Published: (2025)
by: Dobrzycki, Andrzej D., et al.
Published: (2025)
StyleAutoEncoder for manipulating image attributes using pre-trained StyleGAN
by: Bedychaj, Andrzej, et al.
Published: (2024)
by: Bedychaj, Andrzej, et al.
Published: (2024)
TwinMixing: A Shuffle-Aware Feature Interaction Model for Multi-Task Segmentation
by: Do, Minh-Khoi, et al.
Published: (2026)
by: Do, Minh-Khoi, et al.
Published: (2026)
KeyFace: Expressive Audio-Driven Facial Animation for Long Sequences via KeyFrame Interpolation
by: Bigata, Antoni, et al.
Published: (2025)
by: Bigata, Antoni, et al.
Published: (2025)
Cluster and Predict Latent Patches for Improved Masked Image Modeling
by: Darcet, Timothée, et al.
Published: (2025)
by: Darcet, Timothée, et al.
Published: (2025)
Generative Digital Twins: Vision-Language Simulation Models for Executable Industrial Systems
by: Hsu, YuChe, et al.
Published: (2025)
by: Hsu, YuChe, et al.
Published: (2025)
HAWAII: Hierarchical Visual Knowledge Transfer for Efficient Vision-Language Models
by: Wang, Yimu, et al.
Published: (2025)
by: Wang, Yimu, et al.
Published: (2025)
DIAMOND: Directed Inference for Artifact Mitigation in Flow Matching Models
by: Polowczyk, Alicja, et al.
Published: (2026)
by: Polowczyk, Alicja, et al.
Published: (2026)
Weaknesses of Facial Emotion Recognition Systems
by: Jamróz, Aleksandra, et al.
Published: (2026)
by: Jamróz, Aleksandra, et al.
Published: (2026)
Beyond Attention Scores: SVD-Based Vision Token Pruning for Efficient Vision-Language Models
by: Apedo, Yvon, et al.
Published: (2026)
by: Apedo, Yvon, et al.
Published: (2026)
Decoupling the Image Perception and Multimodal Reasoning for Reasoning Segmentation with Digital Twin Representations
by: Li, Yizhen, et al.
Published: (2025)
by: Li, Yizhen, et al.
Published: (2025)
When Pretty Isn't Useful: Investigating Why Modern Text-to-Image Models Fail as Reliable Training Data Generators
by: Adamkiewicz, Krzysztof, et al.
Published: (2026)
by: Adamkiewicz, Krzysztof, et al.
Published: (2026)
Behavior-Grounded Lane Representation Learning for Multi-Task Traffic Digital Twins
by: Tamaru, Rei, et al.
Published: (2026)
by: Tamaru, Rei, et al.
Published: (2026)
A Comprehensive Survey on Surgical Digital Twin
by: Khan, Afsah Sharaf, et al.
Published: (2025)
by: Khan, Afsah Sharaf, et al.
Published: (2025)
Medical Image Segmentation with InTEnt: Integrated Entropy Weighting for Single Image Test-Time Adaptation
by: Dong, Haoyu, et al.
Published: (2024)
by: Dong, Haoyu, et al.
Published: (2024)
Empowering Bridge Digital Twins by Bridging the Data Gap with a Unified Synthesis Framework
by: Wang, Wang, et al.
Published: (2025)
by: Wang, Wang, et al.
Published: (2025)
Self-Evolving Spatial Reasoning in Vision Language Models via Geometric Logic Consistency
by: Liu, Junming, et al.
Published: (2026)
by: Liu, Junming, et al.
Published: (2026)
Entropy Rectifying Guidance for Diffusion and Flow Models
by: Ifriqi, Tariq Berrada, et al.
Published: (2025)
by: Ifriqi, Tariq Berrada, et al.
Published: (2025)
SSL-Interactions: Pretext Tasks for Interactive Trajectory Prediction
by: Bhattacharyya, Prarthana, et al.
Published: (2024)
by: Bhattacharyya, Prarthana, et al.
Published: (2024)
SAR Object Detection with Self-Supervised Pretraining and Curriculum-Aware Sampling
by: Almalioglu, Yasin, et al.
Published: (2025)
by: Almalioglu, Yasin, et al.
Published: (2025)
Facial Width-to-Height Ratio Does Not Predict Self-Reported Behavioral Tendencies
by: Kosinski, Michal
Published: (2024)
by: Kosinski, Michal
Published: (2024)
Geo-ORBIT: A Federated Digital Twin Framework for Scene-Adaptive Lane Geometry Detection
by: Tamaru, Rei, et al.
Published: (2025)
by: Tamaru, Rei, et al.
Published: (2025)
AEGIS: Preserving privacy of 3D Facial Avatars with Adversarial Perturbations
by: Wolkiewicz, Dawid, et al.
Published: (2025)
by: Wolkiewicz, Dawid, et al.
Published: (2025)
The Cost of Language: Centroid Erasure Exposes and Exploits Modal Competition in Multimodal Language Models
by: Paruchuri, Akshay, et al.
Published: (2026)
by: Paruchuri, Akshay, et al.
Published: (2026)
Digital Twin Catalog: A Large-Scale Photorealistic 3D Object Digital Twin Dataset
by: Dong, Zhao, et al.
Published: (2025)
by: Dong, Zhao, et al.
Published: (2025)
Where Do Tokens Go? Understanding Pruning Behaviors in STEP at High Resolutions
by: Szczepanski, Michal, et al.
Published: (2025)
by: Szczepanski, Michal, et al.
Published: (2025)
Less is More: Label-Guided Summarization of Procedural and Instructional Videos
by: Rajpal, Shreya, et al.
Published: (2026)
by: Rajpal, Shreya, et al.
Published: (2026)
Is There Knowledge Left to Extract? Evidence of Fragility in Medically Fine-Tuned Vision-Language Models
by: McLaughlin, Oliver, et al.
Published: (2026)
by: McLaughlin, Oliver, et al.
Published: (2026)
Local Intrinsic Dimension Unveils Hallucinations in Diffusion Models
by: Sobieski, Bartlomiej, et al.
Published: (2026)
by: Sobieski, Bartlomiej, et al.
Published: (2026)
Conditioned Activation Transport for T2I Safety Steering
by: Chrabąszcz, Maciej, et al.
Published: (2026)
by: Chrabąszcz, Maciej, et al.
Published: (2026)
Similar Items
-
Visible Yet Unreadable: A Systematic Blind Spot of Vision Language Models Across Writing Systems
by: Zhang, Jie, et al.
Published: (2025) -
Beyond Recognition: Evaluating Visual Perspective Taking in Vision Language Models
by: Góral, Gracjan, et al.
Published: (2025) -
How Culturally Aware are Vision-Language Models?
by: Burda-Lassen, Olena, et al.
Published: (2024) -
Iconographic Classification and Content-Based Recommendation for Digitized Artworks
by: Kutt, Krzysztof, et al.
Published: (2026) -
A Python toolkit for dealing with Petri nets over ontological graphs
by: Pancerz, Krzysztof
Published: (2025)