A Challenging Benchmark of Anime Style Recognition
Fuente:
arXiv
Guardado en:
| Autores principales: | Li, Haotang, Guo, Shengtao, Lyu, Kailin, Yang, Xiao, Chen, Tianchen, Zhu, Jianqing, Zeng, Huanqiang |
|---|---|
| Formato: | Preprint |
| Publicado: |
2022
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
See-through: Single-image Layer Decomposition for Anime Characters
por: Lin, Jian, et al.
Publicado: (2026)
por: Lin, Jian, et al.
Publicado: (2026)
Positive Style Accumulation: A Style Screening and Continuous Utilization Framework for Federated DG-ReID
por: Xu, Xin, et al.
Publicado: (2025)
por: Xu, Xin, et al.
Publicado: (2025)
U-Net-Like Spiking Neural Networks for Single Image Dehazing
por: Li, Huibin, et al.
Publicado: (2025)
por: Li, Huibin, et al.
Publicado: (2025)
M3CAD: Towards Generic Cooperative Autonomous Driving Benchmark
por: Zhu, Morui, et al.
Publicado: (2025)
por: Zhu, Morui, et al.
Publicado: (2025)
An Evaluation of a Visual Question Answering Strategy for Zero-shot Facial Expression Recognition in Still Images
por: Castrillón-Santana, Modesto, et al.
Publicado: (2025)
por: Castrillón-Santana, Modesto, et al.
Publicado: (2025)
SimWorld: A Unified Benchmark for Simulator-Conditioned Scene Generation via World Model
por: Li, Xinqing, et al.
Publicado: (2025)
por: Li, Xinqing, et al.
Publicado: (2025)
VIAFormer: Voxel-Image Alignment Transformer for High-Fidelity Voxel Refinement
por: Fang, Tiancheng, et al.
Publicado: (2026)
por: Fang, Tiancheng, et al.
Publicado: (2026)
DynFocus: Dynamic Cooperative Network Empowers LLMs with Video Understanding
por: Han, Yudong, et al.
Publicado: (2024)
por: Han, Yudong, et al.
Publicado: (2024)
Context-Aware Network Based on Multi-scale Spatio-temporal Attention for Action Recognition in Videos
por: Li, Xiaoyang, et al.
Publicado: (2025)
por: Li, Xiaoyang, et al.
Publicado: (2025)
Pointing-Based Object Recognition
por: Hajdúch, Lukáš, et al.
Publicado: (2026)
por: Hajdúch, Lukáš, et al.
Publicado: (2026)
MINT: Mitigating Hallucinations in Large Vision-Language Models via Token Reduction
por: Wang, Chao, et al.
Publicado: (2025)
por: Wang, Chao, et al.
Publicado: (2025)
ERNet: Efficient Non-Rigid Registration Network for Point Sequences
por: He, Guangzhao, et al.
Publicado: (2025)
por: He, Guangzhao, et al.
Publicado: (2025)
SurgVLM: A Large Vision-Language Model and Systematic Evaluation Benchmark for Surgical Intelligence
por: Zeng, Zhitao, et al.
Publicado: (2025)
por: Zeng, Zhitao, et al.
Publicado: (2025)
Image-Based Leopard Seal Recognition: Approaches and Challenges in Current Automated Systems
por: Salazar, Jorge Yero, et al.
Publicado: (2024)
por: Salazar, Jorge Yero, et al.
Publicado: (2024)
Synthetic-Child: An AIGC-Based Synthetic Data Pipeline for Privacy-Preserving Child Posture Estimation
por: Zeng, Taowen
Publicado: (2026)
por: Zeng, Taowen
Publicado: (2026)
NAC-TCN: Temporal Convolutional Networks with Causal Dilated Neighborhood Attention for Emotion Understanding
por: Mehta, Alexander, et al.
Publicado: (2023)
por: Mehta, Alexander, et al.
Publicado: (2023)
LEGO: LoRA-Enabled Generator-Oriented Framework for Synthetic Image Detection
por: Xiao, Yutong, et al.
Publicado: (2026)
por: Xiao, Yutong, et al.
Publicado: (2026)
SITransformer: Shared Information-Guided Transformer for Extreme Multimodal Summarization
por: Liu, Sicheng, et al.
Publicado: (2024)
por: Liu, Sicheng, et al.
Publicado: (2024)
GeoHeight-Bench: Towards Height-Aware Multimodal Reasoning in Remote Sensing
por: Hu, Xuran, et al.
Publicado: (2026)
por: Hu, Xuran, et al.
Publicado: (2026)
Lost in Context: The Influence of Context on Feature Attribution Methods for Object Recognition
por: Adhikari, Sayanta, et al.
Publicado: (2024)
por: Adhikari, Sayanta, et al.
Publicado: (2024)
LatentForensics: Towards frugal deepfake detection in the StyleGAN latent space
por: Delmas, Matthieu, et al.
Publicado: (2023)
por: Delmas, Matthieu, et al.
Publicado: (2023)
High-Frequency Semantics and Geometric Priors for End-to-End Detection Transformers in Challenging UAV Imagery
por: Peng, Hongxing, et al.
Publicado: (2025)
por: Peng, Hongxing, et al.
Publicado: (2025)
One Leaf Reveals the Season: Occlusion-Based Contrastive Learning with Semantic-Aware Views for Efficient Visual Representation
por: Yang, Xiaoyu, et al.
Publicado: (2024)
por: Yang, Xiaoyu, et al.
Publicado: (2024)
VisRL: Intention-Driven Visual Perception via Reinforced Reasoning
por: Chen, Zhangquan, et al.
Publicado: (2025)
por: Chen, Zhangquan, et al.
Publicado: (2025)
UniQ: Unified Decoder with Task-specific Queries for Efficient Scene Graph Generation
por: Liao, Xinyao, et al.
Publicado: (2025)
por: Liao, Xinyao, et al.
Publicado: (2025)
Category-Agnostic Neural Object Rigging
por: He, Guangzhao, et al.
Publicado: (2025)
por: He, Guangzhao, et al.
Publicado: (2025)
Visual Enhanced Depth Scaling for Multimodal Latent Reasoning
por: Han, Yudong, et al.
Publicado: (2026)
por: Han, Yudong, et al.
Publicado: (2026)
PlaneSAM: Multimodal Plane Instance Segmentation Using the Segment Anything Model
por: Deng, Zhongchen, et al.
Publicado: (2024)
por: Deng, Zhongchen, et al.
Publicado: (2024)
4DThinker: Thinking with 4D Imagery for Dynamic Spatial Understanding
por: Chen, Zhangquan, et al.
Publicado: (2026)
por: Chen, Zhangquan, et al.
Publicado: (2026)
Cinéaste: A Fine-grained Contextual Movie Question Answering Benchmark
por: Shah, Nisarg A., et al.
Publicado: (2025)
por: Shah, Nisarg A., et al.
Publicado: (2025)
Multi-Scale Spatial-Temporal Self-Attention Graph Convolutional Networks for Skeleton-based Action Recognition
por: Nakamura, Ikuo
Publicado: (2024)
por: Nakamura, Ikuo
Publicado: (2024)
Distinguishing Visually Similar Actions: Prompt-Guided Semantic Prototype Modulation for Few-Shot Action Recognition
por: Li, Xiaoyang, et al.
Publicado: (2025)
por: Li, Xiaoyang, et al.
Publicado: (2025)
NOAH: Benchmarking Narrative Prior driven Hallucination and Omission in Video Large Language Models
por: Lee, Kyuho, et al.
Publicado: (2025)
por: Lee, Kyuho, et al.
Publicado: (2025)
VidNum-1.4K: A Comprehensive Benchmark for Video-based Numerical Reasoning
por: Cui, Shaoyang, et al.
Publicado: (2026)
por: Cui, Shaoyang, et al.
Publicado: (2026)
Domain-Adaptive Pretraining Improves Primate Behavior Recognition
por: Mueller, Felix B., et al.
Publicado: (2025)
por: Mueller, Felix B., et al.
Publicado: (2025)
OCC-MLLM-CoT-Alpha: Towards Multi-stage Occlusion Recognition Based on Large Language Models via 3D-Aware Supervision and Chain-of-Thoughts Guidance
por: Wang, Chaoyi, et al.
Publicado: (2025)
por: Wang, Chaoyi, et al.
Publicado: (2025)
Deep Learning Approaches for Human Action Recognition in Video Data
por: Xie, Yufei
Publicado: (2024)
por: Xie, Yufei
Publicado: (2024)
Application of YOLOv8 in monocular downward multiple Car Target detection
por: Lyu, Shijie
Publicado: (2025)
por: Lyu, Shijie
Publicado: (2025)
EmoVerse: A MLLMs-Driven Emotion Representation Dataset for Interpretable Visual Emotion Analysis
por: Guo, Yijie, et al.
Publicado: (2025)
por: Guo, Yijie, et al.
Publicado: (2025)
Invariant Representation via Decoupling Style and Spurious Features from Images
por: Li, Ruimeng, et al.
Publicado: (2023)
por: Li, Ruimeng, et al.
Publicado: (2023)
Ejemplares similares
-
See-through: Single-image Layer Decomposition for Anime Characters
por: Lin, Jian, et al.
Publicado: (2026) -
Positive Style Accumulation: A Style Screening and Continuous Utilization Framework for Federated DG-ReID
por: Xu, Xin, et al.
Publicado: (2025) -
U-Net-Like Spiking Neural Networks for Single Image Dehazing
por: Li, Huibin, et al.
Publicado: (2025) -
M3CAD: Towards Generic Cooperative Autonomous Driving Benchmark
por: Zhu, Morui, et al.
Publicado: (2025) -
An Evaluation of a Visual Question Answering Strategy for Zero-shot Facial Expression Recognition in Still Images
por: Castrillón-Santana, Modesto, et al.
Publicado: (2025)