Add new CrossEncoder model

Browse files

Files changed (7) hide show

README.md +68 -68
config.json +40 -36
merges.txt +1 -1
onnx/model.onnx +3 -0
special_tokens_map.json +51 -1
tokenizer.json +0 -0
tokenizer_config.json +65 -1

README.md CHANGED Viewed

@@ -1,69 +1,69 @@
----
-language: en
-pipeline_tag: zero-shot-classification
-tags:
-- transformers
-datasets:
-- nyu-mll/multi_nli
-- stanfordnlp/snli
-metrics:
-- accuracy
-license: apache-2.0
-base_model:
-- nreimers/MiniLMv2-L6-H768-distilled-from-RoBERTa-Large
-library_name: sentence-transformers
----
-# Cross-Encoder for Natural Language Inference
-This model was trained using [SentenceTransformers](https://sbert.net) [Cross-Encoder](https://www.sbert.net/examples/applications/cross-encoder/README.html) class.
-## Training Data
-The model was trained on the [SNLI](https://nlp.stanford.edu/projects/snli/) and [MultiNLI](https://cims.nyu.edu/~sbowman/multinli/) datasets. For a given sentence pair, it will output three scores corresponding to the labels: contradiction, entailment, neutral.
-## Performance
-For evaluation results, see [SBERT.net - Pretrained Cross-Encoder](https://www.sbert.net/docs/pretrained_cross-encoders.html#nli).
-## Usage
-Pre-trained models can be used like this:
-```python
-from sentence_transformers import CrossEncoder
-model = CrossEncoder('cross-encoder/nli-MiniLM2-L6-H768')
-scores = model.predict([('A man is eating pizza', 'A man eats something'), ('A black race car starts up in front of a crowd of people.', 'A man is driving down a lonely road.')])
-#Convert scores to labels
-label_mapping = ['contradiction', 'entailment', 'neutral']
-labels = [label_mapping[score_max] for score_max in scores.argmax(axis=1)]
-```
-## Usage with Transformers AutoModel
-You can use the model also directly with Transformers library (without SentenceTransformers library):
-```python
-from transformers import AutoTokenizer, AutoModelForSequenceClassification
-import torch
-model = AutoModelForSequenceClassification.from_pretrained('cross-encoder/nli-MiniLM2-L6-H768')
-tokenizer = AutoTokenizer.from_pretrained('cross-encoder/nli-MiniLM2-L6-H768')
-features = tokenizer(['A man is eating pizza', 'A black race car starts up in front of a crowd of people.'], ['A man eats something', 'A man is driving down a lonely road.'],  padding=True, truncation=True, return_tensors="pt")
-model.eval()
-with torch.no_grad():
-    scores = model(**features).logits
-    label_mapping = ['contradiction', 'entailment', 'neutral']
-    labels = [label_mapping[score_max] for score_max in scores.argmax(dim=1)]
-    print(labels)
-```
-## Zero-Shot Classification
-This model can also be used for zero-shot-classification:
-```python
-from transformers import pipeline
-classifier = pipeline("zero-shot-classification", model='cross-encoder/nli-MiniLM2-L6-H768')
-sent = "Apple just announced the newest iPhone X"
-candidate_labels = ["technology", "sports", "politics"]
-res = classifier(sent, candidate_labels)
-print(res)
 ```

+---
+language: en
+pipeline_tag: zero-shot-classification
+tags:
+- transformers
+datasets:
+- nyu-mll/multi_nli
+- stanfordnlp/snli
+metrics:
+- accuracy
+license: apache-2.0
+base_model:
+- nreimers/MiniLMv2-L6-H768-distilled-from-RoBERTa-Large
+library_name: sentence-transformers
+---
+# Cross-Encoder for Natural Language Inference
+This model was trained using [SentenceTransformers](https://sbert.net) [Cross-Encoder](https://www.sbert.net/examples/applications/cross-encoder/README.html) class.
+## Training Data
+The model was trained on the [SNLI](https://nlp.stanford.edu/projects/snli/) and [MultiNLI](https://cims.nyu.edu/~sbowman/multinli/) datasets. For a given sentence pair, it will output three scores corresponding to the labels: contradiction, entailment, neutral.
+## Performance
+For evaluation results, see [SBERT.net - Pretrained Cross-Encoder](https://www.sbert.net/docs/pretrained_cross-encoders.html#nli).
+## Usage
+Pre-trained models can be used like this:
+```python
+from sentence_transformers import CrossEncoder
+model = CrossEncoder('cross-encoder/nli-MiniLM2-L6-H768')
+scores = model.predict([('A man is eating pizza', 'A man eats something'), ('A black race car starts up in front of a crowd of people.', 'A man is driving down a lonely road.')])
+#Convert scores to labels
+label_mapping = ['contradiction', 'entailment', 'neutral']
+labels = [label_mapping[score_max] for score_max in scores.argmax(axis=1)]
+```
+## Usage with Transformers AutoModel
+You can use the model also directly with Transformers library (without SentenceTransformers library):
+```python
+from transformers import AutoTokenizer, AutoModelForSequenceClassification
+import torch
+model = AutoModelForSequenceClassification.from_pretrained('cross-encoder/nli-MiniLM2-L6-H768')
+tokenizer = AutoTokenizer.from_pretrained('cross-encoder/nli-MiniLM2-L6-H768')
+features = tokenizer(['A man is eating pizza', 'A black race car starts up in front of a crowd of people.'], ['A man eats something', 'A man is driving down a lonely road.'],  padding=True, truncation=True, return_tensors="pt")
+model.eval()
+with torch.no_grad():
+    scores = model(**features).logits
+    label_mapping = ['contradiction', 'entailment', 'neutral']
+    labels = [label_mapping[score_max] for score_max in scores.argmax(dim=1)]
+    print(labels)
+```
+## Zero-Shot Classification
+This model can also be used for zero-shot-classification:
+```python
+from transformers import pipeline
+classifier = pipeline("zero-shot-classification", model='cross-encoder/nli-MiniLM2-L6-H768')
+sent = "Apple just announced the newest iPhone X"
+candidate_labels = ["technology", "sports", "politics"]
+res = classifier(sent, candidate_labels)
+print(res)
 ```

config.json CHANGED Viewed

@@ -1,36 +1,40 @@
-{
-  "_name_or_path": "nreimers/MiniLMv2-L6-H768-distilled-from-RoBERTa-Large",
-  "architectures": [
-    "RobertaForSequenceClassification"
-  ],
-  "attention_probs_dropout_prob": 0.1,
-  "bos_token_id": 0,
-  "eos_token_id": 2,
-  "gradient_checkpointing": false,
-  "hidden_act": "gelu",
-  "hidden_dropout_prob": 0.1,
-  "hidden_size": 768,
-  "id2label": {
-    "0": "contradiction",
-    "1": "entailment",
-    "2": "neutral"
-  },
-  "initializer_range": 0.02,
-  "intermediate_size": 3072,
-  "label2id": {
-    "contradiction": 0,
-    "entailment": 1,
-    "neutral": 2
-  },
-  "layer_norm_eps": 1e-05,
-  "max_position_embeddings": 514,
-  "model_type": "roberta",
-  "num_attention_heads": 12,
-  "num_hidden_layers": 6,
-  "pad_token_id": 1,
-  "position_embedding_type": "absolute",
-  "transformers_version": "4.6.1",
-  "type_vocab_size": 1,
-  "use_cache": true,
-  "vocab_size": 50265
-}

+{
+  "architectures": [
+    "RobertaForSequenceClassification"
+  ],
+  "attention_probs_dropout_prob": 0.1,
+  "bos_token_id": 0,
+  "classifier_dropout": null,
+  "eos_token_id": 2,
+  "gradient_checkpointing": false,
+  "hidden_act": "gelu",
+  "hidden_dropout_prob": 0.1,
+  "hidden_size": 768,
+  "id2label": {
+    "0": "contradiction",
+    "1": "entailment",
+    "2": "neutral"
+  },
+  "initializer_range": 0.02,
+  "intermediate_size": 3072,
+  "label2id": {
+    "contradiction": 0,
+    "entailment": 1,
+    "neutral": 2
+  },
+  "layer_norm_eps": 1e-05,
+  "max_position_embeddings": 514,
+  "model_type": "roberta",
+  "num_attention_heads": 12,
+  "num_hidden_layers": 6,
+  "pad_token_id": 1,
+  "position_embedding_type": "absolute",
+  "sentence_transformers": {
+    "activation_fn": "torch.nn.modules.linear.Identity",
+    "version": "4.1.0.dev0"
+  },
+  "transformers_version": "4.52.0.dev0",
+  "type_vocab_size": 1,
+  "use_cache": true,
+  "vocab_size": 50265
+}

merges.txt CHANGED Viewed

@@ -1,4 +1,4 @@
-#version: 0.2 - Trained by `huggingface/tokenizers`
 Ġ t
 Ġ a
 h e

+#version: 0.2
 Ġ t
 Ġ a
 h e

onnx/model.onnx ADDED Viewed

	@@ -0,0 +1,3 @@

+version https://git-lfs.github.com/spec/v1
+oid sha256:807d33fffefd95ad60e2a91f54eed50a92aa53c077945a6842bb587a49a3fdf3
+size 328649957

special_tokens_map.json CHANGED Viewed

	@@ -1 +1,51 @@
1	- {"bos_token": "<s>", "eos_token": "</s>", "unk_token": "<unk>", "sep_token": "</s>", "pad_token": "<pad>", "cls_token": "<s>", "mask_token": {"content": "<mask>", "single_word": false, "lstrip": true, "rstrip": false, "normalized": false}}

+{
+  "bos_token": {
+    "content": "<s>",
+    "lstrip": false,
+    "normalized": false,
+    "rstrip": false,
+    "single_word": false
+  },
+  "cls_token": {
+    "content": "<s>",
+    "lstrip": false,
+    "normalized": false,
+    "rstrip": false,
+    "single_word": false
+  },
+  "eos_token": {
+    "content": "</s>",
+    "lstrip": false,
+    "normalized": false,
+    "rstrip": false,
+    "single_word": false
+  },
+  "mask_token": {
+    "content": "<mask>",
+    "lstrip": true,
+    "normalized": false,
+    "rstrip": false,
+    "single_word": false
+  },
+  "pad_token": {
+    "content": "<pad>",
+    "lstrip": false,
+    "normalized": false,
+    "rstrip": false,
+    "single_word": false
+  },
+  "sep_token": {
+    "content": "</s>",
+    "lstrip": false,
+    "normalized": false,
+    "rstrip": false,
+    "single_word": false
+  },
+  "unk_token": {
+    "content": "<unk>",
+    "lstrip": false,
+    "normalized": false,
+    "rstrip": false,
+    "single_word": false
+  }
+}

tokenizer.json CHANGED Viewed

The diff for this file is too large to render. See raw diff

tokenizer_config.json CHANGED Viewed

	@@ -1 +1,65 @@
1	- {"unk_token": "<unk>", "bos_token": "<s>", "eos_token": "</s>", "add_prefix_space": false, "errors": "replace", "sep_token": "</s>", "cls_token": "<s>", "pad_token": "<pad>", "mask_token": "<mask>", "model_max_length": 512, "special_tokens_map_file": null, "name_or_path": "nreimers/MiniLMv2-L6-H768-distilled-from-RoBERTa-Large"}

+{
+  "add_prefix_space": false,
+  "added_tokens_decoder": {
+    "0": {
+      "content": "<s>",
+      "lstrip": false,
+      "normalized": false,
+      "rstrip": false,
+      "single_word": false,
+      "special": true
+    },
+    "1": {
+      "content": "<pad>",
+      "lstrip": false,
+      "normalized": false,
+      "rstrip": false,
+      "single_word": false,
+      "special": true
+    },
+    "2": {
+      "content": "</s>",
+      "lstrip": false,
+      "normalized": false,
+      "rstrip": false,
+      "single_word": false,
+      "special": true
+    },
+    "3": {
+      "content": "<unk>",
+      "lstrip": false,
+      "normalized": false,
+      "rstrip": false,
+      "single_word": false,
+      "special": true
+    },
+    "50264": {
+      "content": "<mask>",
+      "lstrip": true,
+      "normalized": false,
+      "rstrip": false,
+      "single_word": false,
+      "special": true
+    }
+  },
+  "bos_token": "<s>",
+  "clean_up_tokenization_spaces": false,
+  "cls_token": "<s>",
+  "eos_token": "</s>",
+  "errors": "replace",
+  "extra_special_tokens": {},
+  "mask_token": "<mask>",
+  "max_length": 512,
+  "model_max_length": 512,
+  "pad_to_multiple_of": null,
+  "pad_token": "<pad>",
+  "pad_token_type_id": 0,
+  "padding_side": "right",
+  "sep_token": "</s>",
+  "stride": 0,
+  "tokenizer_class": "RobertaTokenizer",
+  "trim_offsets": true,
+  "truncation_side": "right",
+  "truncation_strategy": "longest_first",
+  "unk_token": "<unk>"
+}