TexTAR : Textual Attribute Recognition in Multi-domain and Multi-lingual Document Images
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Kumar, Rohan, Jinka, Jyothi Swaroopa, Sarvadevabhatla, Ravi Kiran |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Transfer-LMR: Heavy-Tail Driving Behavior Recognition in Diverse Traffic Scenarios
von: Parikh, Chirag, et al.
Veröffentlicht: (2024)
von: Parikh, Chirag, et al.
Veröffentlicht: (2024)
IndicDLP: A Foundational Dataset for Multi-Lingual and Multi-Domain Document Layout Parsing
von: Nath, Oikantik, et al.
Veröffentlicht: (2025)
von: Nath, Oikantik, et al.
Veröffentlicht: (2025)
MoRAG -- Multi-Fusion Retrieval Augmented Generation for Human Motion
von: Kalakonda, Sai Shashank, et al.
Veröffentlicht: (2024)
von: Kalakonda, Sai Shashank, et al.
Veröffentlicht: (2024)
RoadTones: Tone Controllable Text Generation from Road Event Videos
von: Parikh, Chirag, et al.
Veröffentlicht: (2026)
von: Parikh, Chirag, et al.
Veröffentlicht: (2026)
IDD-X: A Multi-View Dataset for Ego-relative Important Object Localization and Explanation in Dense and Unstructured Traffic
von: Parikh, Chirag, et al.
Veröffentlicht: (2024)
von: Parikh, Chirag, et al.
Veröffentlicht: (2024)
DashCop: Automated E-ticket Generation for Two-Wheeler Traffic Violations Using Dashcam Videos
von: Rawat, Deepti, et al.
Veröffentlicht: (2025)
von: Rawat, Deepti, et al.
Veröffentlicht: (2025)
RoadSocial: A Diverse VideoQA Dataset and Benchmark for Road Event Understanding from Social Video Narratives
von: Parikh, Chirag, et al.
Veröffentlicht: (2025)
von: Parikh, Chirag, et al.
Veröffentlicht: (2025)
SeeThrough3D: Occlusion Aware 3D Control in Text-to-Image Generation
von: Agrawal, Vaibhav, et al.
Veröffentlicht: (2026)
von: Agrawal, Vaibhav, et al.
Veröffentlicht: (2026)
STRinGS: Selective Text Refinement in Gaussian Splatting
von: Raundhal, Abhinav, et al.
Veröffentlicht: (2025)
von: Raundhal, Abhinav, et al.
Veröffentlicht: (2025)
OLAF: A Plug-and-Play Framework for Enhanced Multi-object Multi-part Scene Parsing
von: Gupta, Pranav, et al.
Veröffentlicht: (2024)
von: Gupta, Pranav, et al.
Veröffentlicht: (2024)
CrackUDA: Incremental Unsupervised Domain Adaptation for Improved Crack Segmentation in Civil Structures
von: Srivastava, Kushagra, et al.
Veröffentlicht: (2024)
von: Srivastava, Kushagra, et al.
Veröffentlicht: (2024)
Unveiling Text in Challenging Stone Inscriptions: A Character-Context-Aware Patching Strategy for Binarization
von: Jena, Pratyush, et al.
Veröffentlicht: (2026)
von: Jena, Pratyush, et al.
Veröffentlicht: (2026)
Embedding Textual Information in Images Using Quinary Pixel Combinations
von: Kandala, A V Uday Kiran
Veröffentlicht: (2026)
von: Kandala, A V Uday Kiran
Veröffentlicht: (2026)
MADS: Multi-Attribute Document Supervision for Zero-Shot Image Classification
von: Qu, Xiangyan, et al.
Veröffentlicht: (2025)
von: Qu, Xiangyan, et al.
Veröffentlicht: (2025)
ViTA-PAR: Visual and Textual Attribute Alignment with Attribute Prompting for Pedestrian Attribute Recognition
von: Park, Minjeong, et al.
Veröffentlicht: (2025)
von: Park, Minjeong, et al.
Veröffentlicht: (2025)
TexPainter: Generative Mesh Texturing with Multi-view Consistency
von: Zhang, Hongkun, et al.
Veröffentlicht: (2024)
von: Zhang, Hongkun, et al.
Veröffentlicht: (2024)
MD-ProjTex: Texturing 3D Shapes with Multi-Diffusion Projection
von: Yildirim, Ahmet Burak, et al.
Veröffentlicht: (2025)
von: Yildirim, Ahmet Burak, et al.
Veröffentlicht: (2025)
MixTex: Unambiguous Recognition Should Not Rely Solely on Real Data
von: Luo, Renqing, et al.
Veröffentlicht: (2024)
von: Luo, Renqing, et al.
Veröffentlicht: (2024)
Mask-Guided Multi-Task Network for Face Attribute Recognition
von: Gao, Gong, et al.
Veröffentlicht: (2026)
von: Gao, Gong, et al.
Veröffentlicht: (2026)
Pedestrian Attribute Recognition as Label-balanced Multi-label Learning
von: Zhou, Yibo, et al.
Veröffentlicht: (2024)
von: Zhou, Yibo, et al.
Veröffentlicht: (2024)
BoxingVI: A Multi-Modal Benchmark for Boxing Action Recognition and Localization
von: Kumar, Rahul, et al.
Veröffentlicht: (2025)
von: Kumar, Rahul, et al.
Veröffentlicht: (2025)
FAR-AMTN: Attention Multi-Task Network for Face Attribute Recognition
von: Gao, Gong, et al.
Veröffentlicht: (2026)
von: Gao, Gong, et al.
Veröffentlicht: (2026)
Textual Localization: Decomposing Multi-concept Images for Subject-Driven Text-to-Image Generation
von: Shentu, Junjie, et al.
Veröffentlicht: (2024)
von: Shentu, Junjie, et al.
Veröffentlicht: (2024)
Advancing Textual Prompt Learning with Anchored Attributes
von: Li, Zheng, et al.
Veröffentlicht: (2024)
von: Li, Zheng, et al.
Veröffentlicht: (2024)
Cross-domain Few-shot Object Detection with Multi-modal Textual Enrichment
von: Shangguan, Zeyu, et al.
Veröffentlicht: (2025)
von: Shangguan, Zeyu, et al.
Veröffentlicht: (2025)
TexGen: Text-Guided 3D Texture Generation with Multi-view Sampling and Resampling
von: Huo, Dong, et al.
Veröffentlicht: (2024)
von: Huo, Dong, et al.
Veröffentlicht: (2024)
Unveil: Unified Visual-Textual Integration and Distillation for Multi-modal Document Retrieval
von: Sun, Hao, et al.
Veröffentlicht: (2026)
von: Sun, Hao, et al.
Veröffentlicht: (2026)
SignVTCL: Multi-Modal Continuous Sign Language Recognition Enhanced by Visual-Textual Contrastive Learning
von: Chen, Hao, et al.
Veröffentlicht: (2024)
von: Chen, Hao, et al.
Veröffentlicht: (2024)
Multi Attribute Bias Mitigation via Representation Learning
von: Dwivedi, Rajeev Ranjan, et al.
Veröffentlicht: (2025)
von: Dwivedi, Rajeev Ranjan, et al.
Veröffentlicht: (2025)
EASI-Tex: Edge-Aware Mesh Texturing from Single Image
von: Perla, Sai Raj Kishore, et al.
Veröffentlicht: (2024)
von: Perla, Sai Raj Kishore, et al.
Veröffentlicht: (2024)
Adaptive Prototype Model for Attribute-based Multi-label Few-shot Action Recognition
von: Xiao, Juefeng, et al.
Veröffentlicht: (2025)
von: Xiao, Juefeng, et al.
Veröffentlicht: (2025)
TAR: Text Semantic Assisted Cross-modal Image Registration Framework for Optical and SAR Images
von: Cai, Zhuoyu, et al.
Veröffentlicht: (2026)
von: Cai, Zhuoyu, et al.
Veröffentlicht: (2026)
VISTANet: VIsual Spoken Textual Additive Net for Interpretable Multimodal Emotion Recognition
von: Kumar, Puneet, et al.
Veröffentlicht: (2022)
von: Kumar, Puneet, et al.
Veröffentlicht: (2022)
RomanTex: Decoupling 3D-aware Rotary Positional Embedded Multi-Attention Network for Texture Synthesis
von: Feng, Yifei, et al.
Veröffentlicht: (2025)
von: Feng, Yifei, et al.
Veröffentlicht: (2025)
GenesisTex: Adapting Image Denoising Diffusion to Texture Space
von: Gao, Chenjian, et al.
Veröffentlicht: (2024)
von: Gao, Chenjian, et al.
Veröffentlicht: (2024)
Efficient Multi-domain Text Recognition Deep Neural Network Parameterization with Residual Adapters
von: Chao, Jiayou, et al.
Veröffentlicht: (2024)
von: Chao, Jiayou, et al.
Veröffentlicht: (2024)
ViTAR: Vision Transformer with Any Resolution
von: Fan, Qihang, et al.
Veröffentlicht: (2024)
von: Fan, Qihang, et al.
Veröffentlicht: (2024)
MAKIMA: Tuning-free Multi-Attribute Open-domain Video Editing via Mask-Guided Attention Modulation
von: Zheng, Haoyu, et al.
Veröffentlicht: (2024)
von: Zheng, Haoyu, et al.
Veröffentlicht: (2024)
CustomTex: High-fidelity Indoor Scene Texturing via Multi-Reference Customization
von: Chen, Weilin, et al.
Veröffentlicht: (2026)
von: Chen, Weilin, et al.
Veröffentlicht: (2026)
MSMA: Multi-Scale Feature Fusion For Multi-Attribute 3D Face Reconstruction From Unconstrained Images
von: Cao, Danling
Veröffentlicht: (2025)
von: Cao, Danling
Veröffentlicht: (2025)
Ähnliche Einträge
-
Transfer-LMR: Heavy-Tail Driving Behavior Recognition in Diverse Traffic Scenarios
von: Parikh, Chirag, et al.
Veröffentlicht: (2024) -
IndicDLP: A Foundational Dataset for Multi-Lingual and Multi-Domain Document Layout Parsing
von: Nath, Oikantik, et al.
Veröffentlicht: (2025) -
MoRAG -- Multi-Fusion Retrieval Augmented Generation for Human Motion
von: Kalakonda, Sai Shashank, et al.
Veröffentlicht: (2024) -
RoadTones: Tone Controllable Text Generation from Road Event Videos
von: Parikh, Chirag, et al.
Veröffentlicht: (2026) -
IDD-X: A Multi-View Dataset for Ego-relative Important Object Localization and Explanation in Dense and Unstructured Traffic
von: Parikh, Chirag, et al.
Veröffentlicht: (2024)