SeRpEnt: Selective Resampling for Expressive State Space Models
Fuente:
arXiv
Guardado en:
| Autores principales: | Rando, Stefano, Romani, Luca, Migliarini, Matteo, Franco, Luca, Gudovskiy, Denis, Galasso, Fabio |
|---|---|
| Formato: | Preprint |
| Publicado: |
2025
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Quantifying Self-Preservation Bias in Large Language Models
por: Migliarini, Matteo, et al.
Publicado: (2026)
por: Migliarini, Matteo, et al.
Publicado: (2026)
ReSpace: Text-Driven Autoregressive 3D Indoor Scene Synthesis and Editing
por: Bucher, Martin JJ., et al.
Publicado: (2025)
por: Bucher, Martin JJ., et al.
Publicado: (2025)
HSIMamba: Hyperpsectral Imaging Efficient Feature Learning with Bidirectional State Space for Classification
por: Yang, Judy X, et al.
Publicado: (2024)
por: Yang, Judy X, et al.
Publicado: (2024)
CCS: Clinical Consensus Selection for Radiology Report Generation
por: Zhang, Xi, et al.
Publicado: (2026)
por: Zhang, Xi, et al.
Publicado: (2026)
Predicting 3D Rigid Body Dynamics with Deep Residual Network
por: Oketunji, Abiodun Finbarrs
Publicado: (2024)
por: Oketunji, Abiodun Finbarrs
Publicado: (2024)
Depthwise Separable Convolutions with Deep Residual Convolutions
por: Hasan, Md Arid, et al.
Publicado: (2024)
por: Hasan, Md Arid, et al.
Publicado: (2024)
Content Significance Distribution of Sub-Text Blocks in Articles and Its Application to Article-Organization Assessment
por: Zhou, You, et al.
Publicado: (2023)
por: Zhou, You, et al.
Publicado: (2023)
BabelDOC: Better Layout-Preserving PDF Translation via Intermediate Representation
por: Yang, Qi, et al.
Publicado: (2026)
por: Yang, Qi, et al.
Publicado: (2026)
HalalBench: A Multilingual OCR Benchmark for Food Packaging Ingredient Extraction
por: Arief, Hasan
Publicado: (2026)
por: Arief, Hasan
Publicado: (2026)
R-Genie: Reasoning-Guided Generative Image Editing
por: Zhang, Dong, et al.
Publicado: (2025)
por: Zhang, Dong, et al.
Publicado: (2025)
Cinéaste: A Fine-grained Contextual Movie Question Answering Benchmark
por: Shah, Nisarg A., et al.
Publicado: (2025)
por: Shah, Nisarg A., et al.
Publicado: (2025)
Evaluation Before Generation: A Paradigm for Robust Multimodal Sentiment Analysis with Missing Modalities
por: Chen, Rongfei, et al.
Publicado: (2026)
por: Chen, Rongfei, et al.
Publicado: (2026)
VidNum-1.4K: A Comprehensive Benchmark for Video-based Numerical Reasoning
por: Cui, Shaoyang, et al.
Publicado: (2026)
por: Cui, Shaoyang, et al.
Publicado: (2026)
Proximity QA: Unleashing the Power of Multi-Modal Large Language Models for Spatial Proximity Analysis
por: Li, Jianing, et al.
Publicado: (2024)
por: Li, Jianing, et al.
Publicado: (2024)
Enhanced Kalman with Adaptive Appearance Motion SORT for Grounded Generic Multiple Object Tracking
por: Anh, Duy Le Dinh, et al.
Publicado: (2024)
por: Anh, Duy Le Dinh, et al.
Publicado: (2024)
Multi-Reward GRPO for Stable and Prosodic Single-Codebook TTS LLMs at Scale
por: Zhong, Yicheng, et al.
Publicado: (2025)
por: Zhong, Yicheng, et al.
Publicado: (2025)
Unsupervised Band Selection Using Fused HSI and LiDAR Attention Integrating With Autoencoder
por: Yang, Judy X, et al.
Publicado: (2024)
por: Yang, Judy X, et al.
Publicado: (2024)
On the Expressivity of Selective State-Space Layers: A Multivariate Polynomial Approach
por: Cohen-Karlik, Edo, et al.
Publicado: (2025)
por: Cohen-Karlik, Edo, et al.
Publicado: (2025)
Analyze-Prompt-Reason: A Collaborative Agent-Based Framework for Multi-Image Vision-Language Reasoning
por: Vlachos, Angelos, et al.
Publicado: (2025)
por: Vlachos, Angelos, et al.
Publicado: (2025)
Safeguarding Vision-Language Models Against Patched Visual Prompt Injectors
por: Sun, Jiachen, et al.
Publicado: (2024)
por: Sun, Jiachen, et al.
Publicado: (2024)
Only Whats Necessary: Pareto Optimal Data Minimization for Privacy Preserving Video Anomaly Detection
por: Aslam, Nazia, et al.
Publicado: (2026)
por: Aslam, Nazia, et al.
Publicado: (2026)
SeLeRoSa: Sentence-Level Romanian Satire Detection Dataset
por: Smădu, Răzvan-Alexandru, et al.
Publicado: (2025)
por: Smădu, Răzvan-Alexandru, et al.
Publicado: (2025)
From Pixels to Privacy: Temporally Consistent Video Anonymization via Token Pruning for Privacy Preserving Action Recognition
por: Aslam, Nazia, et al.
Publicado: (2026)
por: Aslam, Nazia, et al.
Publicado: (2026)
Memento 2: Learning by Stateful Reflective Memory
por: Wang, Jun
Publicado: (2025)
por: Wang, Jun
Publicado: (2025)
Towards Blind and Low-Vision Accessibility of Lightweight VLMs and Custom LLM-Evals
por: Baghel, Shruti Singh, et al.
Publicado: (2025)
por: Baghel, Shruti Singh, et al.
Publicado: (2025)
The American Sign Language Knowledge Graph: Infusing ASL Models with Linguistic Knowledge
por: Kezar, Lee, et al.
Publicado: (2024)
por: Kezar, Lee, et al.
Publicado: (2024)
Wildfire spread forecasting with Deep Learning
por: Anastasiou, Nikolaos, et al.
Publicado: (2025)
por: Anastasiou, Nikolaos, et al.
Publicado: (2025)
Beyond Localization: A Comprehensive Diagnosis of Perspective-Conditioned Spatial Reasoning in MLLMs from Omnidirectional Images
por: Chen, Yuangong, et al.
Publicado: (2026)
por: Chen, Yuangong, et al.
Publicado: (2026)
GAEA: A Geolocation Aware Conversational Assistant
por: Campos, Ron, et al.
Publicado: (2025)
por: Campos, Ron, et al.
Publicado: (2025)
SemEval-2025 Task 1: AdMIRe -- Advancing Multimodal Idiomaticity Representation
por: Pickard, Thomas, et al.
Publicado: (2025)
por: Pickard, Thomas, et al.
Publicado: (2025)
GroundCap: A Visually Grounded Image Captioning Dataset
por: Oliveira, Daniel A. P., et al.
Publicado: (2025)
por: Oliveira, Daniel A. P., et al.
Publicado: (2025)
Relative Drawing Identification Complexity is Invariant to Modality in Vision-Language Models
por: Freitas, Diogo, et al.
Publicado: (2025)
por: Freitas, Diogo, et al.
Publicado: (2025)
Fine-Tuning Vision-Language Models for Understanding Current Damage and Scoring Priority with Quality Guard Agent
por: Yasuno, Takato
Publicado: (2026)
por: Yasuno, Takato
Publicado: (2026)
MyoSem: Aligning Electromyography to Natural-Language Action Semantics for Hand Action Understanding
por: Wang, Chiyue, et al.
Publicado: (2026)
por: Wang, Chiyue, et al.
Publicado: (2026)
More Than Meets the Eye: Measuring the Semiotic Gap in Vision-Language Models via Semantic Anchorage
por: He, Wei
Publicado: (2026)
por: He, Wei
Publicado: (2026)
Unpacking the Eye of the Beholder: Social Location, Identity, and the Moving Target of Political Perspectives
por: Sirotkina, Elena
Publicado: (2026)
por: Sirotkina, Elena
Publicado: (2026)
Diffusion Lens: Interpreting Text Encoders in Text-to-Image Pipelines
por: Toker, Michael, et al.
Publicado: (2024)
por: Toker, Michael, et al.
Publicado: (2024)
Story Generation from Visual Inputs: Techniques, Related Tasks, and Challenges
por: Oliveira, Daniel A. P., et al.
Publicado: (2024)
por: Oliveira, Daniel A. P., et al.
Publicado: (2024)
StoryReasoning Dataset: Using Chain-of-Thought for Scene Understanding and Grounded Story Generation
por: Oliveira, Daniel A. P., et al.
Publicado: (2025)
por: Oliveira, Daniel A. P., et al.
Publicado: (2025)
Towards Explainable Fake Image Detection with Multi-Modal Large Language Models
por: Ji, Yikun, et al.
Publicado: (2025)
por: Ji, Yikun, et al.
Publicado: (2025)
Ejemplares similares
-
Quantifying Self-Preservation Bias in Large Language Models
por: Migliarini, Matteo, et al.
Publicado: (2026) -
ReSpace: Text-Driven Autoregressive 3D Indoor Scene Synthesis and Editing
por: Bucher, Martin JJ., et al.
Publicado: (2025) -
HSIMamba: Hyperpsectral Imaging Efficient Feature Learning with Bidirectional State Space for Classification
por: Yang, Judy X, et al.
Publicado: (2024) -
CCS: Clinical Consensus Selection for Radiology Report Generation
por: Zhang, Xi, et al.
Publicado: (2026) -
Predicting 3D Rigid Body Dynamics with Deep Residual Network
por: Oketunji, Abiodun Finbarrs
Publicado: (2024)