vailá: Versatile Anarcho Integrated Liberation Ánalysis in Multimodal Toolbox
Fuente:
arXiv
Guardado en:
| Autores principales: | Santiago, Paulo Roberto Pereira, Chinaglia, Abel Gonçalves, Flanagan, Kira, Bedo, Bruno L. S., Mochida, Ligia Yumi, Aceros, Juan, Bononi, Aline, Cesar, Guilherme Manna |
|---|---|
| Formato: | Preprint |
| Publicado: |
2024
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
PRISM: Differentiable Analysis-by-Synthesis for Fixel Recovery in Diffusion MRI
por: Abouagour, Mohamed, et al.
Publicado: (2026)
por: Abouagour, Mohamed, et al.
Publicado: (2026)
CLIP-Joint-Detect: End-to-End Joint Training of Object Detectors with Contrastive Vision-Language Supervision
por: Raoufi, Behnam, et al.
Publicado: (2025)
por: Raoufi, Behnam, et al.
Publicado: (2025)
Design Patterns for Multilevel Modeling and Simulation
por: Serena, Luca, et al.
Publicado: (2024)
por: Serena, Luca, et al.
Publicado: (2024)
Architecture-Agnostic Feature Synergy for Universal Defense Against Heterogeneous Generative Threats
por: Zhang, Bingxue, et al.
Publicado: (2026)
por: Zhang, Bingxue, et al.
Publicado: (2026)
QoSGMAA: A Robust Multi-Order Graph Attention and Adversarial Framework for Sparse QoS Prediction
por: Du, Guanchen, et al.
Publicado: (2025)
por: Du, Guanchen, et al.
Publicado: (2025)
SALLIE: Safeguarding Against Latent Language & Image Exploits
por: Azov, Guy, et al.
Publicado: (2026)
por: Azov, Guy, et al.
Publicado: (2026)
U-Net-Like Spiking Neural Networks for Single Image Dehazing
por: Li, Huibin, et al.
Publicado: (2025)
por: Li, Huibin, et al.
Publicado: (2025)
The Underlying Dynamics of Life and Its Evolution: A Prigogine-Inspired Informational Dissipative System
por: Chirumbolo, Salvatore, et al.
Publicado: (2024)
por: Chirumbolo, Salvatore, et al.
Publicado: (2024)
Extracting Manifold Information from Point Clouds
por: Guidotti, Patrick
Publicado: (2024)
por: Guidotti, Patrick
Publicado: (2024)
Seeing Is No Longer Believing: Frontier Image Generation Models, Synthetic Visual Evidence, and Real-World Risk
por: Wu, Shuai, et al.
Publicado: (2026)
por: Wu, Shuai, et al.
Publicado: (2026)
Exploring the Capabilities of Large Language Model Encoders for Image-Text Retrieval in Chest X-rays
por: Ko, Hanbin, et al.
Publicado: (2025)
por: Ko, Hanbin, et al.
Publicado: (2025)
Motion Perceiver: Real-Time Occupancy Forecasting for Embedded Systems
por: Ferenczi, Bryce, et al.
Publicado: (2023)
por: Ferenczi, Bryce, et al.
Publicado: (2023)
An integrated heart-torso electromechanical model for the simulation of electrophysiogical outputs accounting for myocardial deformation
por: Zappon, Elena, et al.
Publicado: (2024)
por: Zappon, Elena, et al.
Publicado: (2024)
Pointing-Based Object Recognition
por: Hajdúch, Lukáš, et al.
Publicado: (2026)
por: Hajdúch, Lukáš, et al.
Publicado: (2026)
Learning Association via Track-Detection Matching for Multi-Object Tracking
por: Adžemović, Momir
Publicado: (2025)
por: Adžemović, Momir
Publicado: (2025)
i-DEQ: A stable inertial deep equilibrium model for image restoration
por: Clerc, Antonin, et al.
Publicado: (2026)
por: Clerc, Antonin, et al.
Publicado: (2026)
An Empirical Study for Representations of Videos in Video Question Answering via MLLMs
por: Li, Zhi, et al.
Publicado: (2025)
por: Li, Zhi, et al.
Publicado: (2025)
A Comparative Analysis of Recurrent and Attention Architectures for Isolated Sign Language Recognition
por: Alishzade, Nigar, et al.
Publicado: (2025)
por: Alishzade, Nigar, et al.
Publicado: (2025)
Extracting Explanations, Justification, and Uncertainty from Black-Box Deep Neural Networks
por: Ardis, Paul, et al.
Publicado: (2024)
por: Ardis, Paul, et al.
Publicado: (2024)
Technology prediction of a 3D model using Neural Network
por: Miebs, Grzegorz, et al.
Publicado: (2025)
por: Miebs, Grzegorz, et al.
Publicado: (2025)
High-resolution closed-loop seismic inversion network in time-frequency phase mixed domain
por: Liu, Yingtian, et al.
Publicado: (2024)
por: Liu, Yingtian, et al.
Publicado: (2024)
PhysicsArena: The First Multimodal Physics Reasoning Benchmark Exploring Variable, Process, and Solution Dimensions
por: Dai, Song, et al.
Publicado: (2025)
por: Dai, Song, et al.
Publicado: (2025)
Universal Adversarial Attack on Aligned Multimodal LLMs
por: Rahmatullaev, Temurbek, et al.
Publicado: (2025)
por: Rahmatullaev, Temurbek, et al.
Publicado: (2025)
Game-RL: Synthesizing Multimodal Verifiable Game Data to Boost VLMs' General Reasoning
por: Tong, Jingqi, et al.
Publicado: (2025)
por: Tong, Jingqi, et al.
Publicado: (2025)
SoccerRef-Agents: Multi-Agent System for Automated Soccer Refereeing
por: Meng, Zi, et al.
Publicado: (2026)
por: Meng, Zi, et al.
Publicado: (2026)
AGOP as Explanation: From Feature Learning to Per-Sample Attribution in Image Classifiers
por: Katakam, Raj Kiran Gupta
Publicado: (2026)
por: Katakam, Raj Kiran Gupta
Publicado: (2026)
Memory-Efficient Differentially Private Training with Gradient Random Projection
por: Mulrooney, Alex, et al.
Publicado: (2025)
por: Mulrooney, Alex, et al.
Publicado: (2025)
EduFlow: Advancing MLLMs' Problem-Solving Proficiency through Multi-Stage, Multi-Perspective Critique
por: Zhu, Chenglin, et al.
Publicado: (2025)
por: Zhu, Chenglin, et al.
Publicado: (2025)
PhysNote: Self-Knowledge Notes for Evolvable Physical Reasoning in Vision-Language Model
por: Zhang, Sinin, et al.
Publicado: (2026)
por: Zhang, Sinin, et al.
Publicado: (2026)
Collaborative AI Enhances Image Understanding in Materials Science
por: Yin, Ruoyan Avery, et al.
Publicado: (2025)
por: Yin, Ruoyan Avery, et al.
Publicado: (2025)
ICG: Improving Cover Image Generation via MLLM-based Prompting and Personalized Preference Alignment
por: Bian, Zhipeng, et al.
Publicado: (2026)
por: Bian, Zhipeng, et al.
Publicado: (2026)
OpenMap: Instruction Grounding via Open-Vocabulary Visual-Language Mapping
por: Li, Danyang, et al.
Publicado: (2025)
por: Li, Danyang, et al.
Publicado: (2025)
Dream to Fly: Model-Based Reinforcement Learning for Vision-Based Drone Flight
por: Romero, Angel, et al.
Publicado: (2025)
por: Romero, Angel, et al.
Publicado: (2025)
WayFASTER: a Self-Supervised Traversability Prediction for Increased Navigation Awareness
por: Gasparino, Mateus Valverde, et al.
Publicado: (2024)
por: Gasparino, Mateus Valverde, et al.
Publicado: (2024)
EmbodiedLGR: Integrating Lightweight Graph Representation and Retrieval for Semantic-Spatial Memory in Robotic Agents
por: Riva, Paolo, et al.
Publicado: (2026)
por: Riva, Paolo, et al.
Publicado: (2026)
Taking Flight with Dialogue: Enabling Natural Language Control for PX4-based Drone Agent
por: Lim, Shoon Kit, et al.
Publicado: (2025)
por: Lim, Shoon Kit, et al.
Publicado: (2025)
StratXplore: Strategic Novelty-seeking and Instruction-aligned Exploration for Vision and Language Navigation
por: Gopinathan, Muraleekrishna, et al.
Publicado: (2024)
por: Gopinathan, Muraleekrishna, et al.
Publicado: (2024)
Deep Probabilistic Traversability with Test-time Adaptation for Uncertainty-aware Planetary Rover Navigation
por: Endo, Masafumi, et al.
Publicado: (2024)
por: Endo, Masafumi, et al.
Publicado: (2024)
CoMoCAVs: Cohesive Decision-Guided Motion Planning for Connected and Autonomous Vehicles with Multi-Policy Reinforcement Learning
por: Hu, Pan
Publicado: (2025)
por: Hu, Pan
Publicado: (2025)
Quantized Vision-Language Models for Damage Assessment: A Comparative Study of LLaVA-1.5-7B Quantization Levels
por: Yasuno, Takato
Publicado: (2026)
por: Yasuno, Takato
Publicado: (2026)
Ejemplares similares
-
PRISM: Differentiable Analysis-by-Synthesis for Fixel Recovery in Diffusion MRI
por: Abouagour, Mohamed, et al.
Publicado: (2026) -
CLIP-Joint-Detect: End-to-End Joint Training of Object Detectors with Contrastive Vision-Language Supervision
por: Raoufi, Behnam, et al.
Publicado: (2025) -
Design Patterns for Multilevel Modeling and Simulation
por: Serena, Luca, et al.
Publicado: (2024) -
Architecture-Agnostic Feature Synergy for Universal Defense Against Heterogeneous Generative Threats
por: Zhang, Bingxue, et al.
Publicado: (2026) -
QoSGMAA: A Robust Multi-Order Graph Attention and Adversarial Framework for Sparse QoS Prediction
por: Du, Guanchen, et al.
Publicado: (2025)