lapp0 commited on
Commit
a3488a3
·
verified ·
1 Parent(s): ddd4efe

Training in progress, step 61875

Browse files
README.md CHANGED
@@ -44,42 +44,42 @@ More information needed
44
  | step | epoch | enwikippl | frwikippl | loss | runtime | samples_per_second | steps_per_second | tinystoriesppl | zhwikippl |
45
  | :---: | :---: | :---: | :---: | :---: | :---: | :---: | :---: | :---: | :---: |
46
  | **teacher eval** | | 43.25 | 61.25 | | | | | 11.6875 | 19.125 |
47
- | 0 | 0 | 850403524608.0 | 85212151152640.0 | 21.0960 | 24.9736 | 100.106 | 12.533 | 2952790016.0 | 25013889531904.0 |
48
- | 2500 | 0.0404 | 748.0 | 6560.0 | 2.6187 | 25.0229 | 99.909 | 12.509 | 464.0 | 2656.0 |
49
- | 5000 | 0.0808 | 324.0 | 1416.0 | 1.9058 | 24.9884 | 100.046 | 12.526 | 249.0 | 300.0 |
50
- | 7500 | 0.1212 | 217.0 | 752.0 | 1.6210 | 24.9597 | 100.161 | 12.54 | 181.0 | 190.0 |
51
- | 10000 | 0.1616 | 172.0 | 716.0 | 1.4406 | 25.0292 | 99.883 | 12.505 | 151.0 | 170.0 |
52
- | 12500 | 0.2020 | 124.0 | 458.0 | 1.2049 | 25.0192 | 99.923 | 12.51 | 104.0 | 152.0 |
53
- | 15000 | 0.2424 | 104.5 | 412.0 | 1.0694 | 24.9961 | 100.016 | 12.522 | 88.5 | 145.0 |
54
- | 17500 | 0.2828 | 92.0 | 346.0 | 0.9815 | 25.0156 | 99.937 | 12.512 | 82.0 | 100.0 |
55
- | 20000 | 0.3232 | 83.0 | 314.0 | 0.8990 | 25.0162 | 99.935 | 12.512 | 68.0 | 105.0 |
56
- | 22500 | 0.3636 | 70.5 | 230.0 | 0.7774 | 24.9804 | 100.079 | 12.53 | 57.5 | 72.5 |
57
- | 25000 | 0.4040 | 65.0 | 222.0 | 0.7207 | 25.0102 | 99.959 | 12.515 | 51.75 | 96.5 |
58
- | 27500 | 0.4444 | 64.5 | 202.0 | 0.6891 | 25.0139 | 99.945 | 12.513 | 49.25 | 80.0 |
59
- | 30000 | 0.4848 | 62.0 | 200.0 | 0.6958 | 24.9984 | 100.006 | 12.521 | 48.25 | 66.0 |
60
- | 32500 | 0.5253 | 65.0 | 221.0 | 0.6815 | 25.0096 | 99.962 | 12.515 | 47.5 | 402.0 |
61
- | 35000 | 0.5657 | 60.25 | 189.0 | 0.6250 | 25.0328 | 99.869 | 12.504 | 43.25 | 95.5 |
62
- | 37500 | 0.6061 | 59.0 | 165.0 | 0.6019 | 25.0027 | 99.989 | 12.519 | 43.5 | 72.5 |
63
- | 40000 | 0.6465 | 57.5 | 155.0 | 0.5777 | 24.9823 | 100.071 | 12.529 | 39.0 | 90.5 |
64
- | 42500 | 0.6869 | 59.25 | 161.0 | 0.5669 | 25.0242 | 99.903 | 12.508 | 39.5 | 57.0 |
65
- | 45000 | 0.7273 | 53.0 | 149.0 | 0.4764 | 25.0157 | 99.937 | 12.512 | 34.0 | 42.25 |
66
- | 47500 | 0.7677 | 50.5 | 130.0 | 0.4510 | 24.979 | 100.084 | 12.531 | 33.0 | 40.0 |
67
- | 50000 | 0.8081 | 50.5 | 129.0 | 0.4420 | 24.9761 | 100.096 | 12.532 | 31.75 | 37.0 |
68
- | 52500 | 0.8485 | 50.25 | 134.0 | 0.4333 | 24.9616 | 100.154 | 12.539 | 31.625 | 37.0 |
69
- | 55000 | 0.8889 | 49.25 | 124.0 | 0.4170 | 24.9558 | 100.177 | 12.542 | 30.375 | 35.0 |
70
- | 57500 | 0.9293 | 49.0 | 124.5 | 0.4120 | 24.9149 | 100.342 | 12.563 | 30.25 | 34.75 |
71
- | 60000 | 0.9697 | 49.0 | 123.5 | 0.4089 | 24.9811 | 100.076 | 12.529 | 30.125 | 34.25 |
72
- | 61875 | 1.0 | 49.0 | 123.0 | 0.4086 | 24.9186 | 100.327 | 12.561 | 30.125 | 34.25 |
73
 
74
  # Resource Usage Comparison
75
 
76
- - VRAM Use: 7.7831 GB
77
 
78
- # Distillation (Teacher -> Student) Architecture Difference:
79
 
80
  - **Architecture**: `GPT2LMHeadModel` -> `GPT2LMHeadModel`
81
  - **Total Parameters**: 124,439,808 -> 124,439,808
82
- - **Data Type (dtype)**: torch.bfloat16 -> torch.bfloat16
83
  - **Model Size**: 0.24 GB -> 0.24 GB
84
 
85
  <details>
@@ -93,7 +93,7 @@ More information needed
93
  <br/>
94
 
95
  # Train Dataset
96
- Trained on 145,731,638 tokens from the [wikimedia/wikipedia](https://huggingface.co/datasets/wikimedia/wikipedia) dataset.
97
 
98
  - Num Samples: `247,500`
99
  - Subset: `20231101.en`
@@ -122,7 +122,7 @@ The following hyperparameters were used during training:
122
  - num_epochs: `1.0`
123
  - distillation_objective: `DistillationObjective(logits_loss_component=LossComponent(label=logits, weight=1, loss_fn=kl), attn_loss_component=LossComponent(label=attn, weight=10.0, loss_fn=raw_mse, layer_mapper=layer-2))`
124
  - train_embeddings: `True`
125
- - lr_scheduler: `<torch.optim.lr_scheduler.LambdaLR object at 0x7f6f213d8430>`
126
  - student_model_name_or_path: `None`
127
  - student_config_name_or_path: `None`
128
  - student_model_config: `None`
@@ -154,6 +154,6 @@ The following hyperparameters were used during training:
154
 
155
  # Framework Versions
156
  - Distily 0.2.0
157
- - Transformers 4.44.1
158
- - Pytorch 2.5.0.dev20240821+cu121
159
  - Datasets 2.21.0
 
44
  | step | epoch | enwikippl | frwikippl | loss | runtime | samples_per_second | steps_per_second | tinystoriesppl | zhwikippl |
45
  | :---: | :---: | :---: | :---: | :---: | :---: | :---: | :---: | :---: | :---: |
46
  | **teacher eval** | | 43.25 | 61.25 | | | | | 11.6875 | 19.125 |
47
+ | 0 | 0 | 1159641169920.0 | 74217034874880.0 | 20.1532 | 29.7822 | 83.943 | 10.51 | 3053453312.0 | 65146063945728.0 |
48
+ | 2500 | 0.0404 | 696.0 | 5344.0 | 2.5821 | 29.9253 | 83.541 | 10.459 | 412.0 | 4160.0 |
49
+ | 5000 | 0.0808 | 318.0 | 1336.0 | 1.9054 | 29.8753 | 83.681 | 10.477 | 256.0 | 217.0 |
50
+ | 7500 | 0.1212 | 211.0 | 856.0 | 1.6063 | 29.8573 | 83.732 | 10.483 | 171.0 | 185.0 |
51
+ | 10000 | 0.1616 | 171.0 | 660.0 | 1.4275 | 29.8856 | 83.652 | 10.473 | 137.0 | 151.0 |
52
+ | 12500 | 0.2020 | 120.0 | 454.0 | 1.1681 | 30.0103 | 83.305 | 10.43 | 97.0 | 133.0 |
53
+ | 15000 | 0.2424 | 101.5 | 396.0 | 1.0398 | 29.9702 | 83.416 | 10.444 | 84.0 | 151.0 |
54
+ | 17500 | 0.2828 | 89.0 | 344.0 | 0.9386 | 29.8083 | 83.869 | 10.5 | 71.0 | 130.0 |
55
+ | 20000 | 0.3232 | 80.0 | 296.0 | 0.8577 | 29.9718 | 83.412 | 10.443 | 68.0 | 111.0 |
56
+ | 22500 | 0.3636 | 67.5 | 220.0 | 0.7534 | 29.9265 | 83.538 | 10.459 | 55.75 | 62.5 |
57
+ | 25000 | 0.4040 | 65.0 | 203.0 | 0.7007 | 29.9277 | 83.535 | 10.459 | 51.25 | 121.0 |
58
+ | 27500 | 0.4444 | 64.0 | 212.0 | 0.6707 | 29.8714 | 83.692 | 10.478 | 46.5 | 117.0 |
59
+ | 30000 | 0.4848 | 65.5 | 212.0 | 0.6720 | 29.8654 | 83.709 | 10.48 | 50.25 | 83.5 |
60
+ | 32500 | 0.5253 | 63.5 | 190.0 | 0.6577 | 29.8752 | 83.681 | 10.477 | 47.25 | 57.75 |
61
+ | 35000 | 0.5657 | 60.25 | 176.0 | 0.6085 | 29.8266 | 83.818 | 10.494 | 40.5 | 75.0 |
62
+ | 37500 | 0.6061 | 61.25 | 185.0 | 0.5872 | 29.9842 | 83.377 | 10.439 | 41.25 | 88.5 |
63
+ | 40000 | 0.6465 | 57.75 | 166.0 | 0.5662 | 30.7366 | 81.336 | 10.183 | 41.0 | 64.5 |
64
+ | 42500 | 0.6869 | 58.25 | 173.0 | 0.5596 | 30.0925 | 83.077 | 10.401 | 40.0 | 53.5 |
65
+ | 45000 | 0.7273 | 52.0 | 144.0 | 0.4697 | 29.9028 | 83.604 | 10.467 | 34.75 | 54.0 |
66
+ | 47500 | 0.7677 | 51.75 | 136.0 | 0.4482 | 29.8743 | 83.684 | 10.477 | 33.75 | 40.0 |
67
+ | 50000 | 0.8081 | 50.25 | 142.0 | 0.4373 | 29.8796 | 83.669 | 10.475 | 32.75 | 37.0 |
68
+ | 52500 | 0.8485 | 50.0 | 133.0 | 0.4278 | 29.9384 | 83.505 | 10.455 | 32.5 | 36.0 |
69
+ | 55000 | 0.8889 | 49.0 | 127.5 | 0.4117 | 29.9783 | 83.394 | 10.441 | 31.0 | 35.5 |
70
+ | 57500 | 0.9293 | 49.25 | 129.0 | 0.4067 | 29.9301 | 83.528 | 10.458 | 31.0 | 34.5 |
71
+ | 60000 | 0.9697 | 48.75 | 127.5 | 0.4037 | 29.9613 | 83.441 | 10.447 | 30.875 | 34.25 |
72
+ | 61875 | 1.0 | 49.0 | 127.5 | 0.4034 | 30.7338 | 81.344 | 10.184 | 30.875 | 34.0 |
73
 
74
  # Resource Usage Comparison
75
 
76
+ - VRAM Use: 7.7830 GB
77
 
78
+ `# Distillation (Teacher -> Student) Architecture Difference:
79
 
80
  - **Architecture**: `GPT2LMHeadModel` -> `GPT2LMHeadModel`
81
  - **Total Parameters**: 124,439,808 -> 124,439,808
82
+ - **Data Type (dtype)**: 124439808 -> torch.bfloat16
83
  - **Model Size**: 0.24 GB -> 0.24 GB
84
 
85
  <details>
 
93
  <br/>
94
 
95
  # Train Dataset
96
+ Trained on 145,743,159 tokens from the [wikimedia/wikipedia](https://huggingface.co/datasets/wikimedia/wikipedia) dataset.
97
 
98
  - Num Samples: `247,500`
99
  - Subset: `20231101.en`
 
122
  - num_epochs: `1.0`
123
  - distillation_objective: `DistillationObjective(logits_loss_component=LossComponent(label=logits, weight=1, loss_fn=kl), attn_loss_component=LossComponent(label=attn, weight=10.0, loss_fn=raw_mse, layer_mapper=layer-2))`
124
  - train_embeddings: `True`
125
+ - lr_scheduler: `<torch.optim.lr_scheduler.LambdaLR object at 0x7f01904db0d0>`
126
  - student_model_name_or_path: `None`
127
  - student_config_name_or_path: `None`
128
  - student_model_config: `None`
 
154
 
155
  # Framework Versions
156
  - Distily 0.2.0
157
+ - Transformers 4.44.0
158
+ - Pytorch 2.3.0
159
  - Datasets 2.21.0
config.json CHANGED
@@ -33,7 +33,7 @@
33
  }
34
  },
35
  "torch_dtype": "bfloat16",
36
- "transformers_version": "4.44.1",
37
  "use_cache": true,
38
  "vocab_size": 50257
39
  }
 
33
  }
34
  },
35
  "torch_dtype": "bfloat16",
36
+ "transformers_version": "4.44.0",
37
  "use_cache": true,
38
  "vocab_size": 50257
39
  }
generation_config.json CHANGED
@@ -2,5 +2,5 @@
2
  "_from_model_config": true,
3
  "bos_token_id": 50256,
4
  "eos_token_id": 50256,
5
- "transformers_version": "4.44.1"
6
  }
 
2
  "_from_model_config": true,
3
  "bos_token_id": 50256,
4
  "eos_token_id": 50256,
5
+ "transformers_version": "4.44.0"
6
  }
logs/attn_loss_fn=cos, attn_weight=10.0, projector=ensemble/events.out.tfevents.1724310860.f383272e719b ADDED
@@ -0,0 +1,3 @@
 
 
 
 
1
+ version https://git-lfs.github.com/spec/v1
2
+ oid sha256:b1d236686e30b17f56558c8d281185be7b117961763e91f999aa88e6b3943caa
3
+ size 29632521
logs/attn_loss_fn=raw_mse, attn_weight=10.0, projector=ensemble/completed.flag ADDED
File without changes
model.safetensors CHANGED
@@ -1,3 +1,3 @@
1
  version https://git-lfs.github.com/spec/v1
2
- oid sha256:8f9f176db3a702abc32533c45bbf06a6cae86a50bcd96f10cf37067fdc801b99
3
  size 248894656
 
1
  version https://git-lfs.github.com/spec/v1
2
+ oid sha256:bb261c27a1b53ad2291f783b6d9470d3e6de28aa4c9d0c460f61af0fa157d49e
3
  size 248894656
training_args.bin CHANGED
@@ -1,3 +1,3 @@
1
  version https://git-lfs.github.com/spec/v1
2
- oid sha256:f2685f2af5102dbc593f302ef6e15e3c831f9f625db9d07f72981e7aabd994fe
3
  size 1017899144
 
1
  version https://git-lfs.github.com/spec/v1
2
+ oid sha256:186d249e87d2dcb1f56eba0a694c60e57a50d2f94a427de52c40e818d79b1b4f
3
  size 1017899144