Entity Re-identification in Visual Storytelling via Contrastive Reinforcement Learning
Fuente:
arXiv
Guardado en:
| Autores principales: | Oliveira, Daniel A. P., de Matos, David Martins |
|---|---|
| Formato: | Preprint |
| Publicado: |
2025
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Transfer-learning for video classification: Video Swin Transformer on multiple domains
por: Oliveira, Daniel A. P., et al.
Publicado: (2022)
por: Oliveira, Daniel A. P., et al.
Publicado: (2022)
StoryMovie: A Dataset for Semantic Alignment of Visual Stories with Movie Scripts and Subtitles
por: Oliveira, Daniel, et al.
Publicado: (2026)
por: Oliveira, Daniel, et al.
Publicado: (2026)
Trapped in texture bias? A large scale comparison of deep instance segmentation
por: Theodoridis, Johannes, et al.
Publicado: (2024)
por: Theodoridis, Johannes, et al.
Publicado: (2024)
Enhancing Spatial Reasoning in Vision-Language Models via Chain-of-Thought Prompting and Reinforcement Learning
por: Ji, Binbin, et al.
Publicado: (2025)
por: Ji, Binbin, et al.
Publicado: (2025)
Cytoarchitecture in Words: Weakly Supervised Vision-Language Modeling for Human Brain Microscopy
por: Sutton, Matthew, et al.
Publicado: (2026)
por: Sutton, Matthew, et al.
Publicado: (2026)
Multimodal Approaches for Visually-Rich Document Type Classification: A Comparative Analysis
por: Heyne, Catyana, et al.
Publicado: (2026)
por: Heyne, Catyana, et al.
Publicado: (2026)
Learning from Semantic Dictionaries: Discriminative Codebook Contrastive Learning for Unified Visual Representation and Generation
por: Estepa, Imanol G., et al.
Publicado: (2026)
por: Estepa, Imanol G., et al.
Publicado: (2026)
Sign language recognition based on deep learning and low-cost handcrafted descriptors
por: Carneiro, Alvaro Leandro Cavalcante, et al.
Publicado: (2024)
por: Carneiro, Alvaro Leandro Cavalcante, et al.
Publicado: (2024)
Sparse Concept Bottleneck Models: Gumbel Tricks in Contrastive Learning
por: Semenov, Andrei, et al.
Publicado: (2024)
por: Semenov, Andrei, et al.
Publicado: (2024)
Proximity QA: Unleashing the Power of Multi-Modal Large Language Models for Spatial Proximity Analysis
por: Li, Jianing, et al.
Publicado: (2024)
por: Li, Jianing, et al.
Publicado: (2024)
GAEA: A Geolocation Aware Conversational Assistant
por: Campos, Ron, et al.
Publicado: (2025)
por: Campos, Ron, et al.
Publicado: (2025)
InterChart: Benchmarking Visual Reasoning Across Decomposed and Distributed Chart Information
por: Iyengar, Anirudh Iyengar Kaniyar Narayana, et al.
Publicado: (2025)
por: Iyengar, Anirudh Iyengar Kaniyar Narayana, et al.
Publicado: (2025)
All4One: Symbiotic Neighbour Contrastive Learning via Self-Attention and Redundancy Reduction
por: Estepa, Imanol G., et al.
Publicado: (2023)
por: Estepa, Imanol G., et al.
Publicado: (2023)
Prompt-Driven Building Footprint Extraction in Aerial Images with Offset-Building Model
por: Li, Kai, et al.
Publicado: (2023)
por: Li, Kai, et al.
Publicado: (2023)
GroundCap: A Visually Grounded Image Captioning Dataset
por: Oliveira, Daniel A. P., et al.
Publicado: (2025)
por: Oliveira, Daniel A. P., et al.
Publicado: (2025)
Story Generation from Visual Inputs: Techniques, Related Tasks, and Challenges
por: Oliveira, Daniel A. P., et al.
Publicado: (2024)
por: Oliveira, Daniel A. P., et al.
Publicado: (2024)
MRD: Using Physically Based Differentiable Rendering to Probe Vision Models for 3D Scene Understanding
por: Beilharz, Benjamin, et al.
Publicado: (2025)
por: Beilharz, Benjamin, et al.
Publicado: (2025)
WaveMix: A Resource-efficient Neural Network for Image Analysis
por: Jeevan, Pranav, et al.
Publicado: (2022)
por: Jeevan, Pranav, et al.
Publicado: (2022)
CCVA-FL: Cross-Client Variations Adaptive Federated Learning for Medical Imaging
por: Gupta, Sunny, et al.
Publicado: (2024)
por: Gupta, Sunny, et al.
Publicado: (2024)
Which Backbone to Use: A Resource-efficient Domain Specific Comparison for Computer Vision
por: Jeevan, Pranav, et al.
Publicado: (2024)
por: Jeevan, Pranav, et al.
Publicado: (2024)
Vision transformers in domain adaptation and domain generalization: a study of robustness
por: Alijani, Shadi, et al.
Publicado: (2024)
por: Alijani, Shadi, et al.
Publicado: (2024)
Gaussian Splatting: 3D Reconstruction and Novel View Synthesis, a Review
por: Dalal, Anurag, et al.
Publicado: (2024)
por: Dalal, Anurag, et al.
Publicado: (2024)
StoryReasoning Dataset: Using Chain-of-Thought for Scene Understanding and Grounded Story Generation
por: Oliveira, Daniel A. P., et al.
Publicado: (2025)
por: Oliveira, Daniel A. P., et al.
Publicado: (2025)
FLD+: Data-efficient Evaluation Metric for Generative Models
por: Jeevan, Pranav, et al.
Publicado: (2024)
por: Jeevan, Pranav, et al.
Publicado: (2024)
WaveMixSR-V2: Enhancing Super-resolution with Higher Efficiency
por: Jeevan, Pranav, et al.
Publicado: (2024)
por: Jeevan, Pranav, et al.
Publicado: (2024)
Normalizing Flow-Based Metric for Image Generation
por: Jeevan, Pranav, et al.
Publicado: (2024)
por: Jeevan, Pranav, et al.
Publicado: (2024)
$\textit{sweet}$- An Open Source Modular Platform for Contactless Hand Vascular Biometric Experiments
por: Geissbühler, David, et al.
Publicado: (2024)
por: Geissbühler, David, et al.
Publicado: (2024)
Revisiting Multi-Granularity Representation via Group Contrastive Learning for Unsupervised Vehicle Re-identification
por: Chang, Zhigang, et al.
Publicado: (2024)
por: Chang, Zhigang, et al.
Publicado: (2024)
Canonical Space Representation for 4D Panoptic Segmentation of Articulated Objects
por: Gomes, Manuel, et al.
Publicado: (2025)
por: Gomes, Manuel, et al.
Publicado: (2025)
Fine-Tuning Vision-Language Models for Understanding Current Damage and Scoring Priority with Quality Guard Agent
por: Yasuno, Takato
Publicado: (2026)
por: Yasuno, Takato
Publicado: (2026)
MANGO: Learning Disentangled Image Transformation Manifolds with Grouped Operators
por: Ancelin, Brighton, et al.
Publicado: (2024)
por: Ancelin, Brighton, et al.
Publicado: (2024)
Pan-Arctic Permafrost Landform and Human-built Infrastructure Feature Detection with Vision Transformers and Location Embeddings
por: Perera, Amal S., et al.
Publicado: (2025)
por: Perera, Amal S., et al.
Publicado: (2025)
Large Language Models for Simultaneous Named Entity Extraction and Spelling Correction
por: Whittaker, Edward, et al.
Publicado: (2024)
por: Whittaker, Edward, et al.
Publicado: (2024)
SYNOSIS: Image synthesis pipeline for machine vision in metal surface inspection
por: Fulir, Juraj, et al.
Publicado: (2024)
por: Fulir, Juraj, et al.
Publicado: (2024)
Learnings from Scaling Visual Tokenizers for Reconstruction and Generation
por: Hansen-Estruch, Philippe, et al.
Publicado: (2025)
por: Hansen-Estruch, Philippe, et al.
Publicado: (2025)
Taming the Tail: Leveraging Asymmetric Loss and Pade Approximation to Overcome Medical Image Long-Tailed Class Imbalance
por: Kashyap, Pankhi, et al.
Publicado: (2024)
por: Kashyap, Pankhi, et al.
Publicado: (2024)
Conceptual Evaluation of Deep Visual Stereo Odometry for the MARWIN Radiation Monitoring Robot in Accelerator Tunnels
por: Dehne, André, et al.
Publicado: (2025)
por: Dehne, André, et al.
Publicado: (2025)
Devanagari Handwritten Character Recognition using Convolutional Neural Network
por: Mehta, Diksha, et al.
Publicado: (2025)
por: Mehta, Diksha, et al.
Publicado: (2025)
Low-Cost Tree Crown Dieback Estimation Using Deep Learning-Based Segmentation
por: Allen, M. J., et al.
Publicado: (2024)
por: Allen, M. J., et al.
Publicado: (2024)
Eye-gaze Guided Multi-modal Alignment for Medical Representation Learning
por: Ma, Chong, et al.
Publicado: (2024)
por: Ma, Chong, et al.
Publicado: (2024)
Ejemplares similares
-
Transfer-learning for video classification: Video Swin Transformer on multiple domains
por: Oliveira, Daniel A. P., et al.
Publicado: (2022) -
StoryMovie: A Dataset for Semantic Alignment of Visual Stories with Movie Scripts and Subtitles
por: Oliveira, Daniel, et al.
Publicado: (2026) -
Trapped in texture bias? A large scale comparison of deep instance segmentation
por: Theodoridis, Johannes, et al.
Publicado: (2024) -
Enhancing Spatial Reasoning in Vision-Language Models via Chain-of-Thought Prompting and Reinforcement Learning
por: Ji, Binbin, et al.
Publicado: (2025) -
Cytoarchitecture in Words: Weakly Supervised Vision-Language Modeling for Human Brain Microscopy
por: Sutton, Matthew, et al.
Publicado: (2026)