HyperCLOVA X 32B Think
Fuente:
arXiv
Salvato in:
| Autore principale: | NAVER Cloud HyperCLOVA X Team |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2026
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
HyperCLOVA X 8B Omni
di: NAVER Cloud HyperCLOVA X Team
Pubblicazione: (2026)
di: NAVER Cloud HyperCLOVA X Team
Pubblicazione: (2026)
HyperCLOVA X THINK Technical Report
di: NAVER Cloud HyperCLOVA X Team
Pubblicazione: (2025)
di: NAVER Cloud HyperCLOVA X Team
Pubblicazione: (2025)
HyperCLOVA X Technical Report
di: Yoo, Kang Min, et al.
Pubblicazione: (2024)
di: Yoo, Kang Min, et al.
Pubblicazione: (2024)
Visual Planning: Let's Think Only with Images
di: Xu, Yi, et al.
Pubblicazione: (2025)
di: Xu, Yi, et al.
Pubblicazione: (2025)
Think While You Generate: Discrete Diffusion with Planned Denoising
di: Liu, Sulin, et al.
Pubblicazione: (2024)
di: Liu, Sulin, et al.
Pubblicazione: (2024)
Benchmarking the Thinking Mode of Multimodal Large Language Models in Clinical Tasks
di: Hong, Jindong, et al.
Pubblicazione: (2025)
di: Hong, Jindong, et al.
Pubblicazione: (2025)
Thinking with Patterns: Breaking the Perceptual Bottleneck in Visual Planning via Pattern Induction
di: Jian, Yichang, et al.
Pubblicazione: (2026)
di: Jian, Yichang, et al.
Pubblicazione: (2026)
VisionThink: Smart and Efficient Vision Language Model via Reinforcement Learning
di: Yang, Senqiao, et al.
Pubblicazione: (2025)
di: Yang, Senqiao, et al.
Pubblicazione: (2025)
MMCTAgent: Multi-modal Critical Thinking Agent Framework for Complex Visual Reasoning
di: Kumar, Somnath, et al.
Pubblicazione: (2024)
di: Kumar, Somnath, et al.
Pubblicazione: (2024)
Thinking in Frames: How Visual Context and Test-Time Scaling Empower Video Reasoning
di: Li, Chengzu, et al.
Pubblicazione: (2026)
di: Li, Chengzu, et al.
Pubblicazione: (2026)
xGen-MM-Vid (BLIP-3-Video): You Only Need 32 Tokens to Represent a Video Even in VLMs
di: Ryoo, Michael S., et al.
Pubblicazione: (2024)
di: Ryoo, Michael S., et al.
Pubblicazione: (2024)
LLMs Can Evolve Continually on Modality for X-Modal Reasoning
di: Yu, Jiazuo, et al.
Pubblicazione: (2024)
di: Yu, Jiazuo, et al.
Pubblicazione: (2024)
X-Transfer Attacks: Towards Super Transferable Adversarial Attacks on CLIP
di: Huang, Hanxun, et al.
Pubblicazione: (2025)
di: Huang, Hanxun, et al.
Pubblicazione: (2025)
FluoroSAM: A Language-promptable Foundation Model for Flexible X-ray Image Segmentation
di: Killeen, Benjamin D., et al.
Pubblicazione: (2024)
di: Killeen, Benjamin D., et al.
Pubblicazione: (2024)
SLaVA-CXR: Small Language and Vision Assistant for Chest X-ray Report Automation
di: Wu, Jinge, et al.
Pubblicazione: (2024)
di: Wu, Jinge, et al.
Pubblicazione: (2024)
SPHINX-X: Scaling Data and Parameters for a Family of Multi-modal Large Language Models
di: Liu, Dongyang, et al.
Pubblicazione: (2024)
di: Liu, Dongyang, et al.
Pubblicazione: (2024)
HELPER-X: A Unified Instructable Embodied Agent to Tackle Four Interactive Vision-Language Domains with Memory-Augmented Language Models
di: Sarch, Gabriel, et al.
Pubblicazione: (2024)
di: Sarch, Gabriel, et al.
Pubblicazione: (2024)
CheXpert Plus: Augmenting a Large Chest X-ray Dataset with Text Radiology Reports, Patient Demographics and Additional Image Formats
di: Chambon, Pierre, et al.
Pubblicazione: (2024)
di: Chambon, Pierre, et al.
Pubblicazione: (2024)
CX-Mind: A Pioneering Multimodal Large Language Model for Interleaved Reasoning in Chest X-ray via Curriculum-Guided Reinforcement Learning
di: Li, Wenjie, et al.
Pubblicazione: (2025)
di: Li, Wenjie, et al.
Pubblicazione: (2025)
Physical Property Understanding from Language-Embedded Feature Fields
di: Zhai, Albert J., et al.
Pubblicazione: (2024)
di: Zhai, Albert J., et al.
Pubblicazione: (2024)
PaliGemma: A versatile 3B VLM for transfer
di: Beyer, Lucas, et al.
Pubblicazione: (2024)
di: Beyer, Lucas, et al.
Pubblicazione: (2024)
UI-Venus-1.5 Technical Report
di: Venus Team, et al.
Pubblicazione: (2026)
di: Venus Team, et al.
Pubblicazione: (2026)
When Think-with-Image Meets Safety: What Determines Multimodal Jailbreak Robustness?
di: Tian, Yuan, et al.
Pubblicazione: (2026)
di: Tian, Yuan, et al.
Pubblicazione: (2026)
PAN: A World Model for General, Interactable, and Long-Horizon World Simulation
di: PAN Team, et al.
Pubblicazione: (2025)
di: PAN Team, et al.
Pubblicazione: (2025)
HyperDreamBooth: HyperNetworks for Fast Personalization of Text-to-Image Models
di: Ruiz, Nataniel, et al.
Pubblicazione: (2023)
di: Ruiz, Nataniel, et al.
Pubblicazione: (2023)
Two Experts Are All You Need for Steering Thinking: Reinforcing Cognitive Effort in MoE Reasoning Models Without Additional Training
di: Wang, Mengru, et al.
Pubblicazione: (2025)
di: Wang, Mengru, et al.
Pubblicazione: (2025)
HyperFlow: Gradient-Free Emulation of Few-Shot Fine-Tuning
di: Kim, Donggyun, et al.
Pubblicazione: (2025)
di: Kim, Donggyun, et al.
Pubblicazione: (2025)
HGTS-Former: Hierarchical HyperGraph Transformer for Multivariate Time Series Analysis
di: Si, Hao, et al.
Pubblicazione: (2025)
di: Si, Hao, et al.
Pubblicazione: (2025)
Uncertainty Quantification for Multimodal Large Language Models with Incoherence-adjusted Semantic Volume
di: Lau, Gregory Kang Ruey, et al.
Pubblicazione: (2026)
di: Lau, Gregory Kang Ruey, et al.
Pubblicazione: (2026)
Vision-DeepResearch Benchmark: Rethinking Visual and Textual Search for Multimodal Large Language Models
di: Zeng, Yu, et al.
Pubblicazione: (2026)
di: Zeng, Yu, et al.
Pubblicazione: (2026)
SAVeS: Steering Safety Judgments in Vision-Language Models via Semantic Cues
di: Hinojosa, Carlos, et al.
Pubblicazione: (2026)
di: Hinojosa, Carlos, et al.
Pubblicazione: (2026)
Transformers are Stateless Differentiable Neural Computers
di: Tang, Bo, et al.
Pubblicazione: (2026)
di: Tang, Bo, et al.
Pubblicazione: (2026)
Zooming without Zooming: Region-to-Image Distillation for Fine-Grained Multimodal Perception
di: Wei, Lai, et al.
Pubblicazione: (2026)
di: Wei, Lai, et al.
Pubblicazione: (2026)
A Safety Report on GPT-5.2, Gemini 3 Pro, Qwen3-VL, Grok 4.1 Fast, Nano Banana Pro, and Seedream 4.5
di: Ma, Xingjun, et al.
Pubblicazione: (2026)
di: Ma, Xingjun, et al.
Pubblicazione: (2026)
Video2GUI: Synthesizing Large-Scale Interaction Trajectories for Generalized GUI Agent Pretraining
di: Xiong, Weimin, et al.
Pubblicazione: (2026)
di: Xiong, Weimin, et al.
Pubblicazione: (2026)
When Prompts Override Vision: Prompt-Induced Hallucinations in LVLMs
di: Khayatan, Pegah, et al.
Pubblicazione: (2026)
di: Khayatan, Pegah, et al.
Pubblicazione: (2026)
SpecPL: Disentangling Spectral Granularity for Prompt Learning
di: Zhou, Jingtao, et al.
Pubblicazione: (2026)
di: Zhou, Jingtao, et al.
Pubblicazione: (2026)
Steering the Verifiability of Multimodal AI Hallucinations
di: Pang, Jianhong, et al.
Pubblicazione: (2026)
di: Pang, Jianhong, et al.
Pubblicazione: (2026)
ClawGUI: A Unified Framework for Training, Evaluating, and Deploying GUI Agents
di: Tang, Fei, et al.
Pubblicazione: (2026)
di: Tang, Fei, et al.
Pubblicazione: (2026)
OpenWebRL: Demystifying Online Multi-turn Reinforcement Learning for Visual Web Agents
di: Yang, Rui, et al.
Pubblicazione: (2026)
di: Yang, Rui, et al.
Pubblicazione: (2026)
Documenti analoghi
-
HyperCLOVA X 8B Omni
di: NAVER Cloud HyperCLOVA X Team
Pubblicazione: (2026) -
HyperCLOVA X THINK Technical Report
di: NAVER Cloud HyperCLOVA X Team
Pubblicazione: (2025) -
HyperCLOVA X Technical Report
di: Yoo, Kang Min, et al.
Pubblicazione: (2024) -
Visual Planning: Let's Think Only with Images
di: Xu, Yi, et al.
Pubblicazione: (2025) -
Think While You Generate: Discrete Diffusion with Planned Denoising
di: Liu, Sulin, et al.
Pubblicazione: (2024)