NoTeS-Bank: Benchmarking Neural Transcription and Search for Scientific Notes Understanding
Fuente:
arXiv
Guardado en:
| Autores principales: | Pal, Aniket, Biswas, Sanket, Das, Alloy, Lodh, Ayush, Banerjee, Priyanka, Chattopadhyay, Soumitri, Karatzas, Dimosthenis, Llados, Josep, Jawahar, C. V. |
|---|---|
| Formato: | Preprint |
| Publicado: |
2025
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
FASTER: A Font-Agnostic Scene Text Editing and Rendering Framework
por: Das, Alloy, et al.
Publicado: (2023)
por: Das, Alloy, et al.
Publicado: (2023)
Diving into the Depths of Spotting Text in Multi-Domain Noisy Scenes
por: Das, Alloy, et al.
Publicado: (2023)
por: Das, Alloy, et al.
Publicado: (2023)
Towards Generative Class Prompt Learning for Fine-grained Visual Recognition
por: Chattopadhyay, Soumitri, et al.
Publicado: (2024)
por: Chattopadhyay, Soumitri, et al.
Publicado: (2024)
FastTextSpotter: A High-Efficiency Transformer for Multilingual Scene Text Spotting
por: Das, Alloy, et al.
Publicado: (2024)
por: Das, Alloy, et al.
Publicado: (2024)
Retrieval Augmented Verification for Zero-Shot Detection of Multimodal Disinformation
por: Dey, Arka Ujjal, et al.
Publicado: (2024)
por: Dey, Arka Ujjal, et al.
Publicado: (2024)
GraphKD: Exploring Knowledge Distillation Towards Document Object Detection with Structured Graph Creation
por: Banerjee, Ayan, et al.
Publicado: (2024)
por: Banerjee, Ayan, et al.
Publicado: (2024)
Subjective and Objective Quality Assessment Methods of Stereoscopic Videos with Visibility Affecting Distortions
por: Biswas, Sria, et al.
Publicado: (2024)
por: Biswas, Sria, et al.
Publicado: (2024)
SketchGPT: Autoregressive Modeling for Sketch Generation and Recognition
por: Tiwari, Adarsh, et al.
Publicado: (2024)
por: Tiwari, Adarsh, et al.
Publicado: (2024)
Efficient Transformer-Based Piano Transcription With Sparse Attention Mechanisms
por: Wei, Weixing, et al.
Publicado: (2025)
por: Wei, Weixing, et al.
Publicado: (2025)
HCVR Scene Generation: High Compatibility Virtual Reality Environment Generation for Extended Redirected Walking
por: Zhang, Yiran, et al.
Publicado: (2026)
por: Zhang, Yiran, et al.
Publicado: (2026)
A Distribution Matching Approach to Neural Piano Transcription with Optimal Transport
por: Wei, Weixing, et al.
Publicado: (2026)
por: Wei, Weixing, et al.
Publicado: (2026)
Persistence of Backdoor-based Watermarks for Neural Networks: A Comprehensive Evaluation
por: Ngo, Anh Tu, et al.
Publicado: (2025)
por: Ngo, Anh Tu, et al.
Publicado: (2025)
lifeXplore at the Lifelog Search Challenge 2020
por: Leibetseder, Andreas, et al.
Publicado: (2025)
por: Leibetseder, Andreas, et al.
Publicado: (2025)
TeMTG: Text-Enhanced Multi-Hop Temporal Graph Modeling for Audio-Visual Video Parsing
por: Chen, Yaru, et al.
Publicado: (2025)
por: Chen, Yaru, et al.
Publicado: (2025)
Multimodal LLM-based Query Paraphrasing for Video Search
por: Wu, Jiaxin, et al.
Publicado: (2024)
por: Wu, Jiaxin, et al.
Publicado: (2024)
An Efficient Digital Watermarking Technique for Small Scale devices
por: Talathi, Kaushik, et al.
Publicado: (2025)
por: Talathi, Kaushik, et al.
Publicado: (2025)
CLIPRerank: An Extremely Simple Method for Improving Ad-hoc Video Search
por: Chen, Aozhu, et al.
Publicado: (2024)
por: Chen, Aozhu, et al.
Publicado: (2024)
Harnessing Multimodal Large Language Models for Personalized Product Search with Query-aware Refinement
por: Zhang, Beibei, et al.
Publicado: (2025)
por: Zhang, Beibei, et al.
Publicado: (2025)
PRM-BAS: Enhancing Multimodal Reasoning through PRM-guided Beam Annealing Search
por: Hu, Pengfei, et al.
Publicado: (2025)
por: Hu, Pengfei, et al.
Publicado: (2025)
GeoContrastNet: Contrastive Key-Value Edge Learning for Language-Agnostic Document Understanding
por: Biescas, Nil, et al.
Publicado: (2024)
por: Biescas, Nil, et al.
Publicado: (2024)
MCPNS: A Macropixel Collocated Position and Its Neighbors Search for Plenoptic 2.0 Video Coding
por: Van Duong, Vinh, et al.
Publicado: (2023)
por: Van Duong, Vinh, et al.
Publicado: (2023)
Soundscapes in Spectrograms: Pioneering Multilabel Classification for South Asian Sounds
por: Chakrabarty, Sudip, et al.
Publicado: (2026)
por: Chakrabarty, Sudip, et al.
Publicado: (2026)
Towards Accurate Lip-to-Speech Synthesis in-the-Wild
por: Hegde, Sindhu, et al.
Publicado: (2024)
por: Hegde, Sindhu, et al.
Publicado: (2024)
AGSP-DSA: An Adaptive Graph Signal Processing Framework for Robust Multimodal Fusion with Dynamic Semantic Alignment
por: Karthikeya, KV, et al.
Publicado: (2026)
por: Karthikeya, KV, et al.
Publicado: (2026)
ResearchPulse: Building Method-Experiment Chains through Multi-Document Scientific Inference
por: Chen, Qi, et al.
Publicado: (2025)
por: Chen, Qi, et al.
Publicado: (2025)
mPLUG-PaperOwl: Scientific Diagram Analysis with the Multimodal Large Language Model
por: Hu, Anwen, et al.
Publicado: (2023)
por: Hu, Anwen, et al.
Publicado: (2023)
Variable Rate Image Compression via N-Gram Context based Swin-transformer
por: Mudgal, Priyanka
Publicado: (2025)
por: Mudgal, Priyanka
Publicado: (2025)
DanceAnyWay: Synthesizing Beat-Guided 3D Dances with Randomized Temporal Contrastive Learning
por: Bhattacharya, Aneesh, et al.
Publicado: (2023)
por: Bhattacharya, Aneesh, et al.
Publicado: (2023)
lifeXplore at the Lifelog Search Challenge 2021
por: Leibetseder, Andreas, et al.
Publicado: (2025)
por: Leibetseder, Andreas, et al.
Publicado: (2025)
Unraveling Movie Genres through Cross-Attention Fusion of Bi-Modal Synergy of Poster
por: Nareti, Utsav Kumar, et al.
Publicado: (2024)
por: Nareti, Utsav Kumar, et al.
Publicado: (2024)
Looking Backward: Streaming Video-to-Video Translation with Feature Banks
por: Liang, Feng, et al.
Publicado: (2024)
por: Liang, Feng, et al.
Publicado: (2024)
Unsupervised Transcript-assisted Video Summarization and Highlight Detection
por: Barbakos, Spyros, et al.
Publicado: (2025)
por: Barbakos, Spyros, et al.
Publicado: (2025)
SwinGS: Sliding Window Gaussian Splatting for Volumetric Video Streaming with Arbitrary Length
por: Liu, Bangya, et al.
Publicado: (2024)
por: Liu, Bangya, et al.
Publicado: (2024)
Robust Relevance Feedback for Interactive Known-Item Video Search
por: Ma, Zhixin, et al.
Publicado: (2025)
por: Ma, Zhixin, et al.
Publicado: (2025)
An Efficient Light-weight LSB steganography with Deep learning Steganalysis
por: Das, Dipnarayan, et al.
Publicado: (2022)
por: Das, Dipnarayan, et al.
Publicado: (2022)
Fretting-Transformer: Encoder-Decoder Model for MIDI to Tablature Transcription
por: Hamberger, Anna, et al.
Publicado: (2025)
por: Hamberger, Anna, et al.
Publicado: (2025)
Blurb-Refined Inference from Crowdsourced Book Reviews using Hierarchical Genre Mining with Dual-Path Graph Convolutions
por: Kumar, Suraj, et al.
Publicado: (2025)
por: Kumar, Suraj, et al.
Publicado: (2025)
SlideTailor: Personalized Presentation Slide Generation for Scientific Papers
por: Zeng, Wenzheng, et al.
Publicado: (2025)
por: Zeng, Wenzheng, et al.
Publicado: (2025)
DeepTextMark: A Deep Learning-Driven Text Watermarking Approach for Identifying Large Language Model Generated Text
por: Munyer, Travis, et al.
Publicado: (2023)
por: Munyer, Travis, et al.
Publicado: (2023)
Disentangling Score Content and Performance Style for Joint Piano Rendering and Transcription
por: Zeng, Wei, et al.
Publicado: (2025)
por: Zeng, Wei, et al.
Publicado: (2025)
Ejemplares similares
-
FASTER: A Font-Agnostic Scene Text Editing and Rendering Framework
por: Das, Alloy, et al.
Publicado: (2023) -
Diving into the Depths of Spotting Text in Multi-Domain Noisy Scenes
por: Das, Alloy, et al.
Publicado: (2023) -
Towards Generative Class Prompt Learning for Fine-grained Visual Recognition
por: Chattopadhyay, Soumitri, et al.
Publicado: (2024) -
FastTextSpotter: A High-Efficiency Transformer for Multilingual Scene Text Spotting
por: Das, Alloy, et al.
Publicado: (2024) -
Retrieval Augmented Verification for Zero-Shot Detection of Multimodal Disinformation
por: Dey, Arka Ujjal, et al.
Publicado: (2024)