AIAny
AI Model2026
Icon for item

Pixal3D

Generates high-fidelity 3D assets from a single image by back-projecting pixel-aligned features into 3D, preserving fine geometry and PBR textures; includes inference code and a Hugging Face demo—best suited for single-view object reconstruction.

Introduction

Most single-image 3D methods fuse image features loosely into a 3D backbone, which limits per-pixel fidelity. Pixal3D flips that assumption: it explicitly lifts pixel features into 3D via back-projection to establish direct pixel-to-3D correspondences, enabling near-reconstruction-level detail in both geometry and PBR textures from one view.

Key Capabilities
  • Pixel-aligned lifting: maps 2D pixel features into 3D coordinates rather than relying solely on attention fusion—so what? This preserves fine surface detail and texture alignment that typical implicit or attention-based approaches blur.
  • Single-view reconstruction with PBR textures: produces textured GLB outputs suitable for asset pipelines—so what? You get meshes with material-quality textures that are easier to import into DCC tools and real-time engines.
  • Reproducible stacks & demos: provides branches for the paper implementation and an improved Trellis.2-backed main branch, plus a Hugging Face Gradio demo—so what? You can both reproduce published results and try a higher-performance implementation without local setup.
Who It's For and Trade-offs

Great fit if you need single-image, single-object 3D captures with high fidelity for AR/VR, game assets, or rapid prototyping of product visuals. The project is research-first: it assumes object-centric inputs and depends on a modern backbone (Trellis.2) and nonstandard license—check the repository for commercial terms. Look elsewhere if your target is large scenes, multi-view photogrammetry-grade reconstruction, or extremely low-compute real-time on CPU-only devices, as Pixal3D is optimized around single-view quality and GPU inference.

Where It Fits

Pixal3D sits between generative 3D-from-image models (which prioritize diversity) and reconstruction systems (which prioritize geometric accuracy). Its pixel-to-3D correspondence focus makes it a pragmatic choice when you want higher per-pixel fidelity from a single photograph without full multi-view capture.

Information

  • Websitehuggingface.co
  • AuthorsDong-Yang Li, Wang Zhao, Yuxin Chen, Wenbo Hu, Meng-Hao Guo, Fang-Lue Zhang, Ying Shan, Shi-Min Hu, Tencent ARC Lab
  • Published date2026/04/30

More Items

Hugging Face
AI Model2026

Timestamp-aware realtime video→text model that processes incoming frames continuously, answers questions mid-stream or emits silence when evidence is insufficient, and can revise earlier outputs as new frames arrive. Built for timestamped multimodal interaction with a 256K context and an 11B-parameter backbone.

Hugging Face
AI Model2026

Provides GGUF-quantized Inkling multimodal model weights for local image/audio-to-text and conversational inference. Includes quantization variants (example: 1-bit UD-IQ1_S), Apache-2.0 license, and compatibility with Unsloth Studio, vLLM and common inference stacks.

Hugging Face
AI Video2026

Generates a new camera viewpoint from a reference video: an IC‑LoRA adapter for LTX‑Video 2.3 that re‑renders the same scene from a requested discrete camera angle while preserving subject and content. Trained on synthetic multi‑view data, proof‑of‑concept with limited viewpoint range and best for small, chained angle shifts.