Upload Qwen2ForCausalLM

Files changed (8) hide show

README.md ADDED Viewed

+---
+library_name: transformers
+tags: []
+---
+# GPTQ 4bit quantized version of [DeepSeek-R1-Distill-Qwen-32B](https://huggingface.co/deepseek-ai/DeepSeek-R1-Distill-Qwen-32B)
+## Model Details
+See details in the official model page: [DeepSeek-R1-Distill-Qwen-32B](https://huggingface.co/deepseek-ai/DeepSeek-R1-Distill-Qwen-32B)
+Quantized using [GPTQModel](https://github.com/ModelCloud/GPTQModel) using [wikitext2 dataset](https://github.com/ModelCloud/GPTQModel/blob/main/examples/quantization/basic_usage_wikitext2.py) with `nsamples=512` and `seqlen=2048`. Quantization config:
+```
+bits=4,
+group_size=128,
+desc_act=False,
+damp_percent=0.01,
+```
+Minimum VRAM required: ~20GB
+## How to use
+Using `transformers` library with integrated GPTQ support:
+```python
+from transformers import AutoModelForCausalLM, AutoTokenizer, GenerationConfig
+model_name = "avoroshilov/DeepSeek-R1-Distill-Qwen-32B-GPTQ_4bit-128g"
+tokenizer = AutoTokenizer.from_pretrained(model_name)
+quantized_model = AutoModelForCausalLM.from_pretrained(model_name, device_map='cuda')
+chat = [{"role": "user", "content": "Why is grass green?"},]
+question_tokens = tokenizer.apply_chat_template(chat, add_generation_prompt=True, return_tensors="pt").to(quantized_model.device)
+answer_tokens = quantized_model.generate(question_tokens, generation_config=GenerationConfig(max_length=2048, ))[0]
+print(tokenizer.decode(answer_tokens))
+```

config.json ADDED Viewed

+{
+  "_name_or_path": "DeepSeek-R1-Distill-Qwen-32B-GPTQ_4bit-128g",
+  "architectures": [
+    "Qwen2ForCausalLM"
+  ],
+  "attention_dropout": 0.0,
+  "bos_token_id": 151643,
+  "eos_token_id": 151643,
+  "hidden_act": "silu",
+  "hidden_size": 5120,
+  "initializer_range": 0.02,
+  "intermediate_size": 27648,
+  "max_position_embeddings": 131072,
+  "max_window_layers": 64,
+  "model_type": "qwen2",
+  "num_attention_heads": 40,
+  "num_hidden_layers": 64,
+  "num_key_value_heads": 8,
+  "quantization_config": {
+    "batch_size": 1,
+    "bits": 4,
+    "block_name_to_quantize": null,
+    "cache_block_outputs": true,
+    "damp_percent": 0.1,
+    "dataset": null,
+    "desc_act": false,
+    "exllama_config": {
+      "version": 1
+    },
+    "group_size": 128,
+    "max_input_length": null,
+    "model_seqlen": null,
+    "module_name_preceding_first_block": null,
+    "modules_in_block_to_quantize": null,
+    "pad_token_id": null,
+    "quant_method": "gptq",
+    "sym": true,
+    "tokenizer": null,
+    "true_sequential": true,
+    "use_cuda_fp16": false,
+    "use_exllama": true
+  },
+  "rms_norm_eps": 1e-05,
+  "rope_scaling": null,
+  "rope_theta": 1000000.0,
+  "sliding_window": null,
+  "tie_word_embeddings": false,
+  "torch_dtype": "float16",
+  "transformers_version": "4.48.1",
+  "use_cache": true,
+  "use_sliding_window": false,
+  "vocab_size": 152064
+}

generation_config.json ADDED Viewed

+{
+  "_from_model_config": true,
+  "bos_token_id": 151643,
+  "eos_token_id": 151643,
+  "transformers_version": "4.48.1"
+}

model-00001-of-00004.safetensors ADDED Viewed

+version https://git-lfs.github.com/spec/v1
+oid sha256:7baf591a55c82170d8be3ec7ed570383b8c48305cbc6bb9f9a959429dd8089a0
+size 4960232944

model-00002-of-00004.safetensors ADDED Viewed

+version https://git-lfs.github.com/spec/v1
+oid sha256:a4cc6dd780a37a73954d79d1b307e22bbbd521bd2fae9b935075941104d9ac44
+size 4998127832

model-00003-of-00004.safetensors ADDED Viewed

+version https://git-lfs.github.com/spec/v1
+oid sha256:f0af7df5d1461a161e6bd99145270f31577822b736afcb38fea2c3203052ad28
+size 4965412752

model-00004-of-00004.safetensors ADDED Viewed

+version https://git-lfs.github.com/spec/v1
+oid sha256:03bac45f3e35b7768eb42bd6edab5fb742b568cfe90664e5b660081aa9083d9c
+size 4420211440

model.safetensors.index.json ADDED Viewed

The diff for this file is too large to render. See raw diff