BookNet: Book Image Rectification via Cross-Page Attention Network
Fuente:
arXiv
Saved in:
| Main Authors: | Liu, Shaokai, Feng, Hao, Luan, Bozhi, Hou, Min, Deng, Jiajun, Zhou, Wengang |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
RoFIR: Robust Fisheye Image Rectification Framework Impervious to Optical Center Deviation
by: Liao, Zhaokang, et al.
Published: (2024)
by: Liao, Zhaokang, et al.
Published: (2024)
DeepEraser: Deep Iterative Context Mining for Generic Text Eraser
by: Feng, Hao, et al.
Published: (2024)
by: Feng, Hao, et al.
Published: (2024)
TextCoT: Zoom In for Enhanced Multimodal Text-Rich Image Understanding
by: Luan, Bozhi, et al.
Published: (2024)
by: Luan, Bozhi, et al.
Published: (2024)
Multi-Cue Adaptive Visual Token Pruning for Large Vision-Language Models
by: Luan, Bozhi, et al.
Published: (2025)
by: Luan, Bozhi, et al.
Published: (2025)
ForCenNet: Foreground-Centric Network for Document Image Rectification
by: Cai, Peng, et al.
Published: (2025)
by: Cai, Peng, et al.
Published: (2025)
SwinShadow: Shifted Window for Ambiguous Adjacent Shadow Detection
by: Wang, Yonghui, et al.
Published: (2024)
by: Wang, Yonghui, et al.
Published: (2024)
QueryCDR: Query-Based Controllable Distortion Rectification Network for Fisheye Images
by: Guo, Pengbo, et al.
Published: (2024)
by: Guo, Pengbo, et al.
Published: (2024)
Recurrent Generic Contour-based Instance Segmentation with Progressive Learning
by: Feng, Hao, et al.
Published: (2023)
by: Feng, Hao, et al.
Published: (2023)
DocR1: Evidence Page-Guided GRPO for Multi-Page Document Understanding
by: Xiong, Junyu, et al.
Published: (2025)
by: Xiong, Junyu, et al.
Published: (2025)
Mitigating Object Hallucinations in LVLMs via Attention Imbalance Rectification
by: Sun, Han, et al.
Published: (2026)
by: Sun, Han, et al.
Published: (2026)
CoSMo: A Multimodal Transformer for Page Stream Segmentation in Comic Books
by: Ortega, Marc Serra, et al.
Published: (2025)
by: Ortega, Marc Serra, et al.
Published: (2025)
ShelfRectNet: Single View Shelf Image Rectification with Homography Estimation
by: Tore, Onur Berk, et al.
Published: (2025)
by: Tore, Onur Berk, et al.
Published: (2025)
Self-Classification Enhancement and Correction for Weakly Supervised Object Detection
by: Yin, Yufei, et al.
Published: (2025)
by: Yin, Yufei, et al.
Published: (2025)
Image2Sentence based Asymmetrical Zero-shot Composed Image Retrieval
by: Du, Yongchao, et al.
Published: (2024)
by: Du, Yongchao, et al.
Published: (2024)
A Deep Single Image Rectification Approach for Pan-Tilt-Zoom Cameras
by: Xiao, Teng, et al.
Published: (2025)
by: Xiao, Teng, et al.
Published: (2025)
Revisiting Shadow Detection from a Vision-Language Perspective
by: Wang, Yonghui, et al.
Published: (2026)
by: Wang, Yonghui, et al.
Published: (2026)
AdaptVision: Dynamic Input Scaling in MLLMs for Versatile Scene Understanding
by: Wang, Yonghui, et al.
Published: (2024)
by: Wang, Yonghui, et al.
Published: (2024)
Do You Need Text Rectification? Soft Attention Mask Embedding for Rectification-Free Scene Text Spotting
by: Colombo, Antonio, et al.
Published: (2026)
by: Colombo, Antonio, et al.
Published: (2026)
Y-CA-Net: A Convolutional Attention Based Network for Volumetric Medical Image Segmentation
by: Sharif, Muhammad Hamza, et al.
Published: (2024)
by: Sharif, Muhammad Hamza, et al.
Published: (2024)
Robust Cross-Domain WiFi Fall Detection via Physics-Driven Attention-Enhanced Transformers
by: Wang, Yingzhe, et al.
Published: (2026)
by: Wang, Yingzhe, et al.
Published: (2026)
MoiréNet: A Compact Dual-Domain Network for Image Demoiréing
by: Guo, Shuwei, et al.
Published: (2025)
by: Guo, Shuwei, et al.
Published: (2025)
Parallel Cross Strip Attention Network for Single Image Dehazing
by: Tong, Lihan, et al.
Published: (2024)
by: Tong, Lihan, et al.
Published: (2024)
SMR-Net:Robot Snap Detection Based on Multi-Scale Features and Self-Attention Network
by: Hou, Kuanxu
Published: (2026)
by: Hou, Kuanxu
Published: (2026)
Cascaded Robust Rectification for Arbitrary Document Images
by: Wang, Chaoyun, et al.
Published: (2025)
by: Wang, Chaoyun, et al.
Published: (2025)
nnY-Net: Swin-NeXt with Cross-Attention for 3D Medical Images Segmentation
by: Liu, Haixu, et al.
Published: (2025)
by: Liu, Haixu, et al.
Published: (2025)
Instance-aware Exploration-Verification-Exploitation for Instance ImageGoal Navigation
by: Lei, Xiaohan, et al.
Published: (2024)
by: Lei, Xiaohan, et al.
Published: (2024)
Progressive Multi-modal Conditional Prompt Tuning
by: Qiu, Xiaoyu, et al.
Published: (2024)
by: Qiu, Xiaoyu, et al.
Published: (2024)
StoryDiffusion: Consistent Self-Attention for Long-Range Image and Video Generation
by: Zhou, Yupeng, et al.
Published: (2024)
by: Zhou, Yupeng, et al.
Published: (2024)
Functionalization via Structure Completion and Motion Rectification
by: Zhao, Mingrui, et al.
Published: (2026)
by: Zhao, Mingrui, et al.
Published: (2026)
Layout-to-Image Generation with Localized Descriptions using ControlNet with Cross-Attention Control
by: Lukovnikov, Denis, et al.
Published: (2024)
by: Lukovnikov, Denis, et al.
Published: (2024)
Hierarchical Cross-Attention Network for Virtual Try-On
by: Tang, Hao, et al.
Published: (2024)
by: Tang, Hao, et al.
Published: (2024)
Enhancing Learned Image Compression via Cross Window-based Attention
by: Mudgal, Priyanka, et al.
Published: (2024)
by: Mudgal, Priyanka, et al.
Published: (2024)
MCANet: Medical Image Segmentation with Multi-Scale Cross-Axis Attention
by: Shao, Hao, et al.
Published: (2023)
by: Shao, Hao, et al.
Published: (2023)
ReFrame: Rectification Framework for Image Explaining Architectures
by: Adhikary, Debjyoti Das, et al.
Published: (2025)
by: Adhikary, Debjyoti Das, et al.
Published: (2025)
VoxDepth: Rectification of Depth Images on Edge Devices
by: Chakrabarty, Yashashwee, et al.
Published: (2024)
by: Chakrabarty, Yashashwee, et al.
Published: (2024)
FSCA-Net: Feature-Separated Cross-Attention Network for Robust Multi-Dataset Training
by: Chen, Yuehai
Published: (2026)
by: Chen, Yuehai
Published: (2026)
ScaleWeaver: Weaving Efficient Controllable T2I Generation with Multi-Scale Reference Attention
by: Liu, Keli, et al.
Published: (2025)
by: Liu, Keli, et al.
Published: (2025)
Dual Interaction Network with Cross-Image Attention for Medical Image Segmentation
by: Noh, Jeonghyun, et al.
Published: (2025)
by: Noh, Jeonghyun, et al.
Published: (2025)
GaussNav: Gaussian Splatting for Visual Navigation
by: Lei, Xiaohan, et al.
Published: (2024)
by: Lei, Xiaohan, et al.
Published: (2024)
RAUM-Net: Regional Attention and Uncertainty-aware Mamba Network
by: Liu, Mingquan
Published: (2025)
by: Liu, Mingquan
Published: (2025)
Similar Items
-
RoFIR: Robust Fisheye Image Rectification Framework Impervious to Optical Center Deviation
by: Liao, Zhaokang, et al.
Published: (2024) -
DeepEraser: Deep Iterative Context Mining for Generic Text Eraser
by: Feng, Hao, et al.
Published: (2024) -
TextCoT: Zoom In for Enhanced Multimodal Text-Rich Image Understanding
by: Luan, Bozhi, et al.
Published: (2024) -
Multi-Cue Adaptive Visual Token Pruning for Large Vision-Language Models
by: Luan, Bozhi, et al.
Published: (2025) -
ForCenNet: Foreground-Centric Network for Document Image Rectification
by: Cai, Peng, et al.
Published: (2025)