Upload folder using huggingface_hub

Browse files

Files changed (9) hide show

.gitattributes +2 -0
README.md +63 -3
assets/logo.png +3 -0
assets/teaser.png +3 -0
transformer/config.json +33 -0
transformer/diffusion_pytorch_model-00001-of-00003.safetensors +3 -0
transformer/diffusion_pytorch_model-00002-of-00003.safetensors +3 -0
transformer/diffusion_pytorch_model-00003-of-00003.safetensors +3 -0
transformer/diffusion_pytorch_model.safetensors.index.json +0 -0

.gitattributes CHANGED Viewed

@@ -33,3 +33,5 @@ saved_model/**/* filter=lfs diff=lfs merge=lfs -text
 *.zip filter=lfs diff=lfs merge=lfs -text
 *.zst filter=lfs diff=lfs merge=lfs -text
 *tfevents* filter=lfs diff=lfs merge=lfs -text

 *.zip filter=lfs diff=lfs merge=lfs -text
 *.zst filter=lfs diff=lfs merge=lfs -text
 *tfevents* filter=lfs diff=lfs merge=lfs -text
+assets/logo.png filter=lfs diff=lfs merge=lfs -text
+assets/teaser.png filter=lfs diff=lfs merge=lfs -text

README.md CHANGED Viewed

@@ -1,3 +1,63 @@
----
-license: mit
----

+---
+license: mit
+---
+<div align="center">
+# Aether: Geometric-Aware Unified World Modeling
+</div>
+<div align="center">
+  <img width="400" alt="image" src="assets/logo.png">
+  <!-- <br> -->
+</div>
+<div align="center">
+<a href='https://arxiv.org/abs/2503.18945'><img src='https://img.shields.io/badge/arXiv-2503.18945-red'></a> &nbsp;
+<a href='https://aether-world.github.io'><img src='https://img.shields.io/badge/Project-Page-Green'></a> &nbsp;
+<a href=''><img src='https://img.shields.io/badge/%F0%9F%A4%97%20Hugging%20Face-Demo%20(Coming%20Soon)-blue'></a> &nbsp;
+</div>
+Aether addresses a fundamental challenge in AI: integrating geometric reconstruction with generative modeling
+for human-like spatial reasoning. Our framework unifies three core capabilities: (1) **4D dynamic reconstruction**,
+(2) **action-conditioned video prediction**, and (3) **goal-conditioned visual planning**. Trained entirely on
+synthetic data, Aether achieves strong zero-shot generalization to real-world scenarios.
+<div align="center">
+    <img src="assets/teaser.png" alt="Teaser" width="800"/>
+</div>
+## 📝 Citation
+If you find this work useful in your research, please consider citing:
+```bibtex
+@article{aether,
+  title     = {Aether: Geometric-Aware Unified World Modeling},
+  author    = {Aether Team and Haoyi Zhu and Yifan Wang and Jianjun Zhou and Wenzheng Chang and Yang Zhou and Zizun Li and Junyi Chen and Chunhua Shen and Jiangmiao Pang and Tong He},
+  journal   = {arXiv preprint arXiv:2503.18945},
+  year      = {2025}
+}
+```
+## ⚖️ License
+This repository is licensed under the MIT License - see the [LICENSE](LICENSE) file for details.
+## 🙏 Acknowledgements
+Our work is primarily built upon
+[Accelerate](https://github.com/huggingface/accelerate),
+[Diffusers](https://github.com/huggingface/diffusers),
+[CogVideoX](https://github.com/THUDM/CogVideo),
+[Finetrainers](https://github.com/a-r-r-o-w/finetrainers),
+[DepthAnyVideo](https://github.com/Nightmare-n/DepthAnyVideo),
+[CUT3R](https://github.com/CUT3R/CUT3R),
+[MonST3R](https://github.com/Junyi42/monst3r),
+[VBench](https://github.com/Vchitect/VBench),
+[GST](https://github.com/SOTAMak1r/GST),
+[SPA](https://github.com/HaoyiZhu/SPA),
+[DroidCalib](https://github.com/boschresearch/DroidCalib),
+[Grounded-SAM-2](https://github.com/IDEA-Research/Grounded-SAM-2),
+[ceres-solver](https://github.com/ceres-solver/ceres-solver), etc.
+We extend our gratitude to all these authors for their generously open-sourced code and their significant contributions to the community.

assets/logo.png ADDED Viewed

Git LFS Details

SHA256: 1fcc6a3c8e5fc8206ce96ca50f85b06aa337d38354b98b4faef986f06026550e
Pointer size: 132 Bytes
Size of remote file: 1.29 MB

assets/teaser.png ADDED Viewed

Git LFS Details

SHA256: 3b9cfe7dbabbb999ad75f78ef3a38ffb0ed9f56303cff3a0d9ebfa90bf29031c
Pointer size: 133 Bytes
Size of remote file: 11.8 MB

transformer/config.json ADDED Viewed

	@@ -0,0 +1,33 @@

+{
+  "_class_name": "CogVideoXTransformer3DModel",
+  "_diffusers_version": "0.32.2",
+  "_name_or_path": "THUDM/CogVideoX-5b-I2V",
+  "activation_fn": "gelu-approximate",
+  "attention_bias": true,
+  "attention_head_dim": 64,
+  "dropout": 0.0,
+  "flip_sin_to_cos": true,
+  "freq_shift": 0,
+  "in_channels": 96,
+  "max_text_seq_length": 226,
+  "norm_elementwise_affine": true,
+  "norm_eps": 1e-05,
+  "num_attention_heads": 48,
+  "num_layers": 42,
+  "ofs_embed_dim": null,
+  "out_channels": 56,
+  "patch_bias": true,
+  "patch_size": 2,
+  "patch_size_t": null,
+  "sample_frames": 41,
+  "sample_height": 60,
+  "sample_width": 90,
+  "spatial_interpolation_scale": 1.875,
+  "temporal_compression_ratio": 4,
+  "temporal_interpolation_scale": 1.0,
+  "text_embed_dim": 4096,
+  "time_embed_dim": 512,
+  "timestep_activation_fn": "silu",
+  "use_learned_positional_embeddings": false,
+  "use_rotary_positional_embeddings": true
+}

transformer/diffusion_pytorch_model-00001-of-00003.safetensors ADDED Viewed

	@@ -0,0 +1,3 @@

+version https://git-lfs.github.com/spec/v1
+oid sha256:ba6d3d89bb92a2d9c42e025090317477b2e653b6e081c61d311b6aff866ef020
+size 4979268296

transformer/diffusion_pytorch_model-00002-of-00003.safetensors ADDED Viewed

	@@ -0,0 +1,3 @@

+version https://git-lfs.github.com/spec/v1
+oid sha256:69db4f65a4e99f0ff7fc05574287d3264fb7c1114edfd108d921a89c58640b4e
+size 4948039832

transformer/diffusion_pytorch_model-00003-of-00003.safetensors ADDED Viewed

	@@ -0,0 +1,3 @@

+version https://git-lfs.github.com/spec/v1
+oid sha256:52691657f95be290feb229f0665cb11c97fa47a479ec2b9c44e6cb94a3f4b20c
+size 1216323744

transformer/diffusion_pytorch_model.safetensors.index.json ADDED Viewed

The diff for this file is too large to render. See raw diff