prithivMLmods commited on
Commit
a42acdf
·
verified ·
1 Parent(s): 31eb87c

Update README.md

Browse files
Files changed (1) hide show
  1. README.md +9 -5
README.md CHANGED
@@ -9,22 +9,26 @@ tags:
9
  - merge
10
 
11
  ---
12
- # merge
 
 
 
 
 
13
 
14
  This is a merge of pre-trained language models created using [mergekit](https://github.com/cg123/mergekit).
15
 
16
- ## Merge Details
17
- ### Merge Method
18
 
19
  This model was merged using the [TIES](https://arxiv.org/abs/2306.01708) merge method using [Qwen/Qwen2.5-7B-Instruct-1M](https://huggingface.co/Qwen/Qwen2.5-7B-Instruct-1M) as a base.
20
 
21
- ### Models Merged
22
 
23
  The following models were included in the merge:
24
  * [deepseek-ai/DeepSeek-R1-Distill-Qwen-7B](https://huggingface.co/deepseek-ai/DeepSeek-R1-Distill-Qwen-7B)
25
  * [Qwen/Qwen2.5-7B-Instruct](https://huggingface.co/Qwen/Qwen2.5-7B-Instruct)
26
 
27
- ### Configuration
28
 
29
  The following YAML configuration was used to produce this model:
30
 
 
9
  - merge
10
 
11
  ---
12
+
13
+
14
+ # **Qwen2.5-7B-DeepSeek-R1-1M**
15
+
16
+ This model is a merged pre-trained language model created using MergeKit with the TIES merge method. It uses **Qwen/Qwen2.5-7B-Instruct-1M** as the base and combines **deepseek-ai/DeepSeek-R1-Distill-Qwen-7B** and **Qwen/Qwen2.5-7B-Instruct** with equal weight and density. The merge configuration includes normalization, int8 masking, and `bfloat16` precision for optimized performance.
17
+ # **Merge**
18
 
19
  This is a merge of pre-trained language models created using [mergekit](https://github.com/cg123/mergekit).
20
 
21
+ # **Merge Method**
 
22
 
23
  This model was merged using the [TIES](https://arxiv.org/abs/2306.01708) merge method using [Qwen/Qwen2.5-7B-Instruct-1M](https://huggingface.co/Qwen/Qwen2.5-7B-Instruct-1M) as a base.
24
 
25
+ # **Models Merged**
26
 
27
  The following models were included in the merge:
28
  * [deepseek-ai/DeepSeek-R1-Distill-Qwen-7B](https://huggingface.co/deepseek-ai/DeepSeek-R1-Distill-Qwen-7B)
29
  * [Qwen/Qwen2.5-7B-Instruct](https://huggingface.co/Qwen/Qwen2.5-7B-Instruct)
30
 
31
+ # **Configuration**
32
 
33
  The following YAML configuration was used to produce this model:
34