Add new SentenceTransformer model

Browse files

Files changed (11) hide show

1_Pooling/config.json +10 -0
README.md +613 -0
config.json +25 -0
config_sentence_transformers.json +12 -0
model.safetensors +3 -0
modules.json +20 -0
sentence_bert_config.json +4 -0
special_tokens_map.json +37 -0
tokenizer.json +0 -0
tokenizer_config.json +63 -0
vocab.txt +0 -0

1_Pooling/config.json ADDED Viewed

	@@ -0,0 +1,10 @@

+{
+  "word_embedding_dimension": 1024,
+  "pooling_mode_cls_token": true,
+  "pooling_mode_mean_tokens": false,
+  "pooling_mode_max_tokens": false,
+  "pooling_mode_mean_sqrt_len_tokens": false,
+  "pooling_mode_weightedmean_tokens": false,
+  "pooling_mode_lasttoken": false,
+  "include_prompt": true
+}

README.md ADDED Viewed

	@@ -0,0 +1,613 @@

+---
+tags:
+- sentence-transformers
+- sentence-similarity
+- feature-extraction
+- generated_from_trainer
+- dataset_size:400
+- loss:MatryoshkaLoss
+- loss:MultipleNegativesRankingLoss
+base_model: Snowflake/snowflake-arctic-embed-l
+widget:
+- source_sentence: What actions did Mr and Mrs Harris take that led to the revelation
+    of the facts in the case?
+  sentences:
+  - "Perplexity’s marketing activities include promoting on its Instagram account\
+    \ a massive billboard \nin Times Square from September 2024 which read “Congratulations\
+    \ Perplexity on 250 million \nquestions answered last month.”5 \n \n4 Discover\
+    \ New York with Perplexity, Perplexity AI (last visited Oct. 17, 2024), \nhttps://www.perplexity.ai/encyclopedia/discovernewyork.\
+    \ \n5 @perplexity.ai, Instagram (Sept. 4, 2024),  \nhttps://www.instagram.com/perplexity.ai/p/C_g2TonSHC5.\
+    \  \nCase 1:24-cv-07984     Document 1     Filed 10/21/24     Page 8 of 42"
+  - "31 \n \nstatus.    It was not until Mr. and Mrs. Harris retained counsel, served\
+    \ a demand letter on May 22, \n2024, met with the then Assistant Superintendent\
+    \ and a lengthy “bulling investigation” that these \nfacts came to light.    \n\
+    The Defendant’s actions and conduct, by definition, was arbitrary and capricious\
+    \ as was \nthe  imposition of discipline that was a gross abuse of discretion\
+    \ when it served as a catalyst for \nthis action.  Similarly, the Defendants exceeded\
+    \ their authority by repeatedly doubling down on \ntheir acts and conduct when\
+    \ given the opportunity to reverse course.  The adverse action taken was \nnot\
+    \ based on sound, objective, adopted and approved policies and procedures regarding\
+    \ the use of"
+  - "website users, and licensing is transacted with individuals and entities residing\
+    \ in this State and \nDistrict. As such, the injuries alleged herein from Perplexity’s\
+    \ infringement and other unlawful \nconduct foreseeably occurred in this State\
+    \ and District. In addition, Perplexity or its agents reside \nin this District\
+    \ and may be found in this State and District. \n23. \nDefendant Perplexity is\
+    \ subject to the jurisdiction of this Court pursuant to N.Y. \nC.P.L.R. § 302(a)(1)\
+    \ and (3) as it has purposefully directed its activities at New York and has \n\
+    Case 1:24-cv-07984     Document 1     Filed 10/21/24     Page 7 of 42"
+- source_sentence: How did the Plaintiffs demonstrate that the discipline and sanctions
+    imposed by Hingham were arbitrary and capricious?
+  sentences:
+  - "27 \n \nin the adoption and execution of policies and practices that in their\
+    \ judgment are needed to preserve \ninternal order and discipline and to maintain\
+    \ institutional security\"), such deference is not without \nlimitation.  The\
+    \ propriety of and deference afforded to the decision making is a rebuttable \n\
+    presumption that may only be undone by a showing that the action taken was arbitrary\
+    \ and \ncapricious.  See Doe v. Supt. Of Schools of Stoughton, 437 Mass. 1, 5\
+    \ (2002).  The Plaintiffs are \nlikely to succeed on the merits because they have\
+    \ shown through Hingham’s own investigation \nmaterials that the discipline and\
+    \ sanctions imposed were arbitrary, capricious and an abuse of \ndiscretion under\
+    \ the circumstances."
+  - "7 \n \nhighly competitive curriculum with by and large top grades, a 36 ACT (highest\
+    \ score possible) \nand a varied solid resume.  In order for RNH to apply to Stanford\
+    \ by November 1, which means \nsubmitting no later than October 25th, his transcript\
+    \ issue must be resolved by early October so \nthat when RNH requests his transcripts,\
+    \ they reflect grades commensurate with his achievement \nand not marred by the\
+    \ incident that gave rise to this case. \nLetter grades of “C” in this type of\
+    \ admissions environment typically lead to the applicant \nbeing excluded from\
+    \ consideration.  Additionally, transcripts and information regarding any \ndisciplinary\
+    \ infraction, especially one regarding an academic integrity infraction, are a\
+    \ substantial"
+  - "as one’s own.  Id. at ¶106-107.  During the project, RNH and his classmate did\
+    \ not take someone \nelse’s work or ideas and pass them off as their own.  Id.\
+    \ at ¶108.  RNH and his classmate used AI, \nwhich generates and synthesizes new\
+    \ information, and did not pass off another’s work as their \nown.  Id. at ¶109.\
+    \  Despite having this information, the Defendants exceeded the authority granted\
+    \ \nto them in an abuse of authority, discretion, and unfettered state action\
+    \ by unfairly and unjustly \nacting as investigator, judge, jury, and executioner\
+    \ in determining the extreme and outrageous \nsanctions imposed upon these Students.\
+    \  Id. at ¶110.   \nAfter being unfairly and unjustly accused of cheating, plagiarism,\
+    \ and academic"
+- source_sentence: How many students with academic infractions were inducted into
+    the NHS, and what was one of the reasons for their infractions?
+  sentences:
+  - "companies that want to utilize popular, high-quality, human-created journalism\
+    \ for use by the \ncompanies’ AI applications. Revenue received from legitimately-run\
+    \ AI companies supports the \ncosts of news gathering. This revenue also establishes\
+    \ that there is a market for the licensing of \nhuman-generated content for lawful\
+    \ use in AI technologies. \n42. \nPlaintiffs’ content is highly valued in this\
+    \ market.  \n \n \nCase 1:24-cv-07984     Document 1     Filed 10/21/24     Page\
+    \ 12 of 42"
+  - "18 \n \nupon in affirming the decision through an appeal to exclude RNH and his\
+    \ classmate from the NHS.  \nId. at ¶145.  At that time, Defendant Swanson and\
+    \ other Defendants knew or should have known \nthat the District inducted at least\
+    \ seven students into NHS, who had academic infractions on their \nrecord, one\
+    \ of which was because of the prior use of AI.  Id. at ¶146.   \nThe “committee”\
+    \ that adjudicated selection for NHS this year did not include teachers who \n\
+    know and are familiar with RNH and his classmate.  Id. at ¶147.  This is due to\
+    \ the then escalating \ncontract conflict with the Hingham Educators Association\
+    \ (“HEA”) where HEA engaged in an"
+  - "is plainly and solely for a commercial purpose. Moreover, upon information and\
+    \ belief, it copies \ninto its index every single word of Plaintiffs’ copyrighted\
+    \ works that it can get its hands on. \nAdditionally, the use to which it puts\
+    \ these copies is to create a commercial substitute for Plaintiffs’ \nprotected\
+    \ works – in Perplexity’s own words, to allow and encourage users to “Skip the\
+    \ Links” to \nPlaintiffs’ original works. Such substitution causes substantial\
+    \ harm to Plaintiffs’ traditional \nadvertising and subscription revenues. Perplexity’s\
+    \ conduct also harms Plaintiffs’ additional, \nestablished revenue stream from\
+    \ licensing to more scrupulous AI companies. Nor is Perplexity’s"
+- source_sentence: How many pages does the document filed in case 1:24-cv-12437-WGY
+    contain?
+  sentences:
+  - Case 1:24-cv-12437-WGY   Document 8   Filed 10/08/24   Page 26 of 42
+  - "more specialized in conducting a specific task, responding to prompts specific\
+    \ to a subject area, or \nrecognizing nuances in particular questions. \n48. \n\
+    Fine-tuning might also ensure that the LLM responds to certain prompts by \nmimicking\
+    \ a certain linguistic style. For example, outputting a cooking recipe requires\
+    \ a distinct \noutput style from recounting the statistics of the Allies’ landing\
+    \ at Normandy’s beaches on D-Day, \nor from writing a poem about the summer wind.\
+    \ A medical treatise uses a distinct linguistic style \nfrom a sports recap. \n\
+    Case 1:24-cv-07984     Document 1     Filed 10/21/24     Page 13 of 42"
+  - "D. The Balancing Of The Irreparable Harm Heavily Favors The Plaintiffs \nThe\
+    \ balance of harms in this case clearly favors granting an injunction. If the\
+    \ injunction is \nnot granted, RNH will suffer irreparable harm that cannot be\
+    \ adequately remedied by any future \ncourt decision or monetary compensation.\
+    \ RNH’s academic and professional future is at stake, as \na delayed resolution\
+    \ of the investigation into academic sanctions could result in missed deadlines\
+    \ \nfor college applications, exclusion from consideration at elite universities,\
+    \ and a permanent stain \non his academic record. The reputational damage and\
+    \ uncertainty caused will undermine RNH’s \nability to compete fairly with other\
+    \ applicants, affecting not only his immediate educational"
+- source_sentence: What challenges do professional journalists and publishers face
+    that may impact their ability to enforce their intellectual property rights?
+  sentences:
+  - "ban or prohibition on the use of AI by students. The Defendants were not trained\
+    \ on any policies \nor procedures for use of AI alone, never mind what they were\
+    \ “able to do” to students who used \nit.    The entire purpose behind having\
+    \ such policies and procedures in place is to ensure notice, \nequity, fairness\
+    \ and to be sure:  a level playing field for all.   Making matters worse, there\
+    \ exists \nno adequate procedures and policies for the induction of an applicant\
+    \ into NHS when compared to \nother members who are inducted despite the same\
+    \ or similar infractions.  This is a denial of student \nrights of the highest\
+    \ order. \n \nIn the case here, RNH was disciplined on an ad hoc and on-going\
+    \ basis over more than six"
+  - "19 \nrespect. They feel very good about it. And in our user interface, even though\
+    \ we give the answer, \nwe do show the user exactly where the answer is coming\
+    \ from.”16  \n68. \nAs Srinivas surely knows or should know, academic standards\
+    \ for avoiding \nplagiarism are wholly independent from copyright law.17 Dow Jones\
+    \ and NYP Holdings editors \nand journalists are not graduate students working\
+    \ out of a library or lab, eager to have someone \nacknowledge and utilize their\
+    \ research. They are professional journalists and publishers – working \nunder\
+    \ high-pressure deadlines, sometimes in dangerous places – whose livelihoods depend\
+    \ on the \nenforcement and monetization of their intellectual property rights.\
+    \  \n69."
+  - "example the school committee under Mass. G.L. c. 71, § 37, may punish a student\
+    \ offender without \na prior rule specifically forbidding the offending conduct;\
+    \ however, surely such authority cannot \nbe limitless. Moreover, this court believes\
+    \ that the imposition of a severe penalty without a \nspecific promulgated rule\
+    \ might be constitutionally deficient under certain circumstances.   \nId. (emphasis\
+    \ supplied). “What those circumstances are can only be left to the development\
+    \ of the \ncase law in the area.”    Id.   There has been no case law developed\
+    \ in the area of school discipline \nCase 1:24-cv-12437-WGY   Document 8   Filed\
+    \ 10/08/24   Page 28 of 42"
+pipeline_tag: sentence-similarity
+library_name: sentence-transformers
+metrics:
+- cosine_accuracy@1
+- cosine_accuracy@3
+- cosine_accuracy@5
+- cosine_accuracy@10
+- cosine_precision@1
+- cosine_precision@3
+- cosine_precision@5
+- cosine_precision@10
+- cosine_recall@1
+- cosine_recall@3
+- cosine_recall@5
+- cosine_recall@10
+- cosine_ndcg@10
+- cosine_mrr@10
+- cosine_map@100
+model-index:
+- name: SentenceTransformer based on Snowflake/snowflake-arctic-embed-l
+  results:
+  - task:
+      type: information-retrieval
+      name: Information Retrieval
+    dataset:
+      name: Unknown
+      type: unknown
+    metrics:
+    - type: cosine_accuracy@1
+      value: 0.7291666666666666
+      name: Cosine Accuracy@1
+    - type: cosine_accuracy@3
+      value: 0.8541666666666666
+      name: Cosine Accuracy@3
+    - type: cosine_accuracy@5
+      value: 0.9375
+      name: Cosine Accuracy@5
+    - type: cosine_accuracy@10
+      value: 1.0
+      name: Cosine Accuracy@10
+    - type: cosine_precision@1
+      value: 0.7291666666666666
+      name: Cosine Precision@1
+    - type: cosine_precision@3
+      value: 0.28472222222222215
+      name: Cosine Precision@3
+    - type: cosine_precision@5
+      value: 0.1875
+      name: Cosine Precision@5
+    - type: cosine_precision@10
+      value: 0.09999999999999999
+      name: Cosine Precision@10
+    - type: cosine_recall@1
+      value: 0.7291666666666666
+      name: Cosine Recall@1
+    - type: cosine_recall@3
+      value: 0.8541666666666666
+      name: Cosine Recall@3
+    - type: cosine_recall@5
+      value: 0.9375
+      name: Cosine Recall@5
+    - type: cosine_recall@10
+      value: 1.0
+      name: Cosine Recall@10
+    - type: cosine_ndcg@10
+      value: 0.8575788154610162
+      name: Cosine Ndcg@10
+    - type: cosine_mrr@10
+      value: 0.8125248015873017
+      name: Cosine Mrr@10
+    - type: cosine_map@100
+      value: 0.8125248015873016
+      name: Cosine Map@100
+---
+# SentenceTransformer based on Snowflake/snowflake-arctic-embed-l
+This is a [sentence-transformers](https://www.SBERT.net) model finetuned from [Snowflake/snowflake-arctic-embed-l](https://huggingface.co/Snowflake/snowflake-arctic-embed-l). It maps sentences & paragraphs to a 1024-dimensional dense vector space and can be used for semantic textual similarity, semantic search, paraphrase mining, text classification, clustering, and more.
+## Model Details
+### Model Description
+- **Model Type:** Sentence Transformer
+- **Base model:** [Snowflake/snowflake-arctic-embed-l](https://huggingface.co/Snowflake/snowflake-arctic-embed-l) <!-- at revision d8fb21ca8d905d2832ee8b96c894d3298964346b -->
+- **Maximum Sequence Length:** 512 tokens
+- **Output Dimensionality:** 1024 dimensions
+- **Similarity Function:** Cosine Similarity
+<!-- - **Training Dataset:** Unknown -->
+<!-- - **Language:** Unknown -->
+<!-- - **License:** Unknown -->
+### Model Sources
+- **Documentation:** [Sentence Transformers Documentation](https://sbert.net)
+- **Repository:** [Sentence Transformers on GitHub](https://github.com/UKPLab/sentence-transformers)
+- **Hugging Face:** [Sentence Transformers on Hugging Face](https://huggingface.co/models?library=sentence-transformers)
+### Full Model Architecture
+```
+SentenceTransformer(
+  (0): Transformer({'max_seq_length': 512, 'do_lower_case': False}) with Transformer model: BertModel
+  (1): Pooling({'word_embedding_dimension': 1024, 'pooling_mode_cls_token': True, 'pooling_mode_mean_tokens': False, 'pooling_mode_max_tokens': False, 'pooling_mode_mean_sqrt_len_tokens': False, 'pooling_mode_weightedmean_tokens': False, 'pooling_mode_lasttoken': False, 'include_prompt': True})
+  (2): Normalize()
+)
+```
+## Usage
+### Direct Usage (Sentence Transformers)
+First install the Sentence Transformers library:
+```bash
+pip install -U sentence-transformers
+```
+Then you can load this model and run inference.
+```python
+from sentence_transformers import SentenceTransformer
+# Download from the 🤗 Hub
+model = SentenceTransformer("llm-wizard/legal-ft")
+# Run inference
+sentences = [
+    'What challenges do professional journalists and publishers face that may impact their ability to enforce their intellectual property rights?',
+    '19 \nrespect. They feel very good about it. And in our user interface, even though we give the answer, \nwe do show the user exactly where the answer is coming from.”16  \n68. \nAs Srinivas surely knows or should know, academic standards for avoiding \nplagiarism are wholly independent from copyright law.17 Dow Jones and NYP Holdings editors \nand journalists are not graduate students working out of a library or lab, eager to have someone \nacknowledge and utilize their research. They are professional journalists and publishers – working \nunder high-pressure deadlines, sometimes in dangerous places – whose livelihoods depend on the \nenforcement and monetization of their intellectual property rights.  \n69.',
+    'ban or prohibition on the use of AI by students. The Defendants were not trained on any policies \nor procedures for use of AI alone, never mind what they were “able to do” to students who used \nit.    The entire purpose behind having such policies and procedures in place is to ensure notice, \nequity, fairness and to be sure:  a level playing field for all.   Making matters worse, there exists \nno adequate procedures and policies for the induction of an applicant into NHS when compared to \nother members who are inducted despite the same or similar infractions.  This is a denial of student \nrights of the highest order. \n \nIn the case here, RNH was disciplined on an ad hoc and on-going basis over more than six',
+]
+embeddings = model.encode(sentences)
+print(embeddings.shape)
+# [3, 1024]
+# Get the similarity scores for the embeddings
+similarities = model.similarity(embeddings, embeddings)
+print(similarities.shape)
+# [3, 3]
+```
+<!--
+### Direct Usage (Transformers)
+<details><summary>Click to see the direct usage in Transformers</summary>
+</details>
+-->
+<!--
+### Downstream Usage (Sentence Transformers)
+You can finetune this model on your own dataset.
+<details><summary>Click to expand</summary>
+</details>
+-->
+<!--
+### Out-of-Scope Use
+*List how the model may foreseeably be misused and address what users ought not to do with the model.*
+-->
+## Evaluation
+### Metrics
+#### Information Retrieval
+* Evaluated with [<code>InformationRetrievalEvaluator</code>](https://sbert.net/docs/package_reference/sentence_transformer/evaluation.html#sentence_transformers.evaluation.InformationRetrievalEvaluator)
+| Metric              | Value      |
+|:--------------------|:-----------|
+| cosine_accuracy@1   | 0.7292     |
+| cosine_accuracy@3   | 0.8542     |
+| cosine_accuracy@5   | 0.9375     |
+| cosine_accuracy@10  | 1.0        |
+| cosine_precision@1  | 0.7292     |
+| cosine_precision@3  | 0.2847     |
+| cosine_precision@5  | 0.1875     |
+| cosine_precision@10 | 0.1        |
+| cosine_recall@1     | 0.7292     |
+| cosine_recall@3     | 0.8542     |
+| cosine_recall@5     | 0.9375     |
+| cosine_recall@10    | 1.0        |
+| **cosine_ndcg@10**  | **0.8576** |
+| cosine_mrr@10       | 0.8125     |
+| cosine_map@100      | 0.8125     |
+<!--
+## Bias, Risks and Limitations
+*What are the known or foreseeable issues stemming from this model? You could also flag here known failure cases or weaknesses of the model.*
+-->
+<!--
+### Recommendations
+*What are recommendations with respect to the foreseeable issues? For example, filtering explicit content.*
+-->
+## Training Details
+### Training Dataset
+#### Unnamed Dataset
+* Size: 400 training samples
+* Columns: <code>sentence_0</code> and <code>sentence_1</code>
+* Approximate statistics based on the first 400 samples:
+  |         | sentence_0                                                                         | sentence_1                                                                           |
+  |:--------|:-----------------------------------------------------------------------------------|:-------------------------------------------------------------------------------------|
+  | type    | string                                                                             | string                                                                               |
+  | details | <ul><li>min: 10 tokens</li><li>mean: 20.93 tokens</li><li>max: 35 tokens</li></ul> | <ul><li>min: 25 tokens</li><li>mean: 140.37 tokens</li><li>max: 260 tokens</li></ul> |
+* Samples:
+  | sentence_0                                                                                                             | sentence_1                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                            |
+  |:-----------------------------------------------------------------------------------------------------------------------|:--------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------|
+  | <code>What provisions of the 2023-2024 Handbook were referenced regarding the use of AI and academic integrity?</code> | <code>13 <br> <br>procedure, expectation, conduct, discipline, sanction or consequence for the use of AI.  Id. at ¶102.  <br>Under these circumstances, the use of AI was not a violation of the then existing “Academic <br>Integrity: Cheating and Plagiarism” provisions of the 2023-2024 Handbook.  Id. at ¶104.  As such, <br>accusations of cheating, plagiarism, and academic misconduct or dishonesty were not supported <br>by the record evidence which, at all times relevant, the Defendants have had in their care, custody <br>and control.  Id. at ¶105.   <br>While there is much dispute as to whether the use of generative AI constitutes plagiarism, <br>plagiarism is defined as the practice of taking someone else’s work or ideas and passing them off</code> |
+  | <code>How is plagiarism defined in the context provided?</code>                                                        | <code>13 <br> <br>procedure, expectation, conduct, discipline, sanction or consequence for the use of AI.  Id. at ¶102.  <br>Under these circumstances, the use of AI was not a violation of the then existing “Academic <br>Integrity: Cheating and Plagiarism” provisions of the 2023-2024 Handbook.  Id. at ¶104.  As such, <br>accusations of cheating, plagiarism, and academic misconduct or dishonesty were not supported <br>by the record evidence which, at all times relevant, the Defendants have had in their care, custody <br>and control.  Id. at ¶105.   <br>While there is much dispute as to whether the use of generative AI constitutes plagiarism, <br>plagiarism is defined as the practice of taking someone else’s work or ideas and passing them off</code> |
+  | <code>What is the case number associated with the document filed on 10/21/24?</code>                                   | <code>program-ad-revenue-sharing-ai-time-fortune-der-spiegel. <br>Case 1:24-cv-07984     Document 1     Filed 10/21/24     Page 21 of 42</code>                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                       |
+* Loss: [<code>MatryoshkaLoss</code>](https://sbert.net/docs/package_reference/sentence_transformer/losses.html#matryoshkaloss) with these parameters:
+  ```json
+  {
+      "loss": "MultipleNegativesRankingLoss",
+      "matryoshka_dims": [
+          768,
+          512,
+          256,
+          128,
+          64
+      ],
+      "matryoshka_weights": [
+          1,
+          1,
+          1,
+          1,
+          1
+      ],
+      "n_dims_per_step": -1
+  }
+  ```
+### Training Hyperparameters
+#### Non-Default Hyperparameters
+- `eval_strategy`: steps
+- `per_device_train_batch_size`: 10
+- `per_device_eval_batch_size`: 10
+- `num_train_epochs`: 10
+- `multi_dataset_batch_sampler`: round_robin
+#### All Hyperparameters
+<details><summary>Click to expand</summary>
+- `overwrite_output_dir`: False
+- `do_predict`: False
+- `eval_strategy`: steps
+- `prediction_loss_only`: True
+- `per_device_train_batch_size`: 10
+- `per_device_eval_batch_size`: 10
+- `per_gpu_train_batch_size`: None
+- `per_gpu_eval_batch_size`: None
+- `gradient_accumulation_steps`: 1
+- `eval_accumulation_steps`: None
+- `torch_empty_cache_steps`: None
+- `learning_rate`: 5e-05
+- `weight_decay`: 0.0
+- `adam_beta1`: 0.9
+- `adam_beta2`: 0.999
+- `adam_epsilon`: 1e-08
+- `max_grad_norm`: 1
+- `num_train_epochs`: 10
+- `max_steps`: -1
+- `lr_scheduler_type`: linear
+- `lr_scheduler_kwargs`: {}
+- `warmup_ratio`: 0.0
+- `warmup_steps`: 0
+- `log_level`: passive
+- `log_level_replica`: warning
+- `log_on_each_node`: True
+- `logging_nan_inf_filter`: True
+- `save_safetensors`: True
+- `save_on_each_node`: False
+- `save_only_model`: False
+- `restore_callback_states_from_checkpoint`: False
+- `no_cuda`: False
+- `use_cpu`: False
+- `use_mps_device`: False
+- `seed`: 42
+- `data_seed`: None
+- `jit_mode_eval`: False
+- `use_ipex`: False
+- `bf16`: False
+- `fp16`: False
+- `fp16_opt_level`: O1
+- `half_precision_backend`: auto
+- `bf16_full_eval`: False
+- `fp16_full_eval`: False
+- `tf32`: None
+- `local_rank`: 0
+- `ddp_backend`: None
+- `tpu_num_cores`: None
+- `tpu_metrics_debug`: False
+- `debug`: []
+- `dataloader_drop_last`: False
+- `dataloader_num_workers`: 0
+- `dataloader_prefetch_factor`: None
+- `past_index`: -1
+- `disable_tqdm`: False
+- `remove_unused_columns`: True
+- `label_names`: None
+- `load_best_model_at_end`: False
+- `ignore_data_skip`: False
+- `fsdp`: []
+- `fsdp_min_num_params`: 0
+- `fsdp_config`: {'min_num_params': 0, 'xla': False, 'xla_fsdp_v2': False, 'xla_fsdp_grad_ckpt': False}
+- `fsdp_transformer_layer_cls_to_wrap`: None
+- `accelerator_config`: {'split_batches': False, 'dispatch_batches': None, 'even_batches': True, 'use_seedable_sampler': True, 'non_blocking': False, 'gradient_accumulation_kwargs': None}
+- `deepspeed`: None
+- `label_smoothing_factor`: 0.0
+- `optim`: adamw_torch
+- `optim_args`: None
+- `adafactor`: False
+- `group_by_length`: False
+- `length_column_name`: length
+- `ddp_find_unused_parameters`: None
+- `ddp_bucket_cap_mb`: None
+- `ddp_broadcast_buffers`: False
+- `dataloader_pin_memory`: True
+- `dataloader_persistent_workers`: False
+- `skip_memory_metrics`: True
+- `use_legacy_prediction_loop`: False
+- `push_to_hub`: False
+- `resume_from_checkpoint`: None
+- `hub_model_id`: None
+- `hub_strategy`: every_save
+- `hub_private_repo`: None
+- `hub_always_push`: False
+- `gradient_checkpointing`: False
+- `gradient_checkpointing_kwargs`: None
+- `include_inputs_for_metrics`: False
+- `include_for_metrics`: []
+- `eval_do_concat_batches`: True
+- `fp16_backend`: auto
+- `push_to_hub_model_id`: None
+- `push_to_hub_organization`: None
+- `mp_parameters`:
+- `auto_find_batch_size`: False
+- `full_determinism`: False
+- `torchdynamo`: None
+- `ray_scope`: last
+- `ddp_timeout`: 1800
+- `torch_compile`: False
+- `torch_compile_backend`: None
+- `torch_compile_mode`: None
+- `dispatch_batches`: None
+- `split_batches`: None
+- `include_tokens_per_second`: False
+- `include_num_input_tokens_seen`: False
+- `neftune_noise_alpha`: None
+- `optim_target_modules`: None
+- `batch_eval_metrics`: False
+- `eval_on_start`: False
+- `use_liger_kernel`: False
+- `eval_use_gather_object`: False
+- `average_tokens_across_devices`: False
+- `prompts`: None
+- `batch_sampler`: batch_sampler
+- `multi_dataset_batch_sampler`: round_robin
+</details>
+### Training Logs
+| Epoch | Step | cosine_ndcg@10 |
+|:-----:|:----:|:--------------:|
+| 1.0   | 40   | 0.8182         |
+| 1.25  | 50   | 0.8172         |
+| 2.0   | 80   | 0.8112         |
+| 2.5   | 100  | 0.8414         |
+| 3.0   | 120  | 0.8236         |
+| 3.75  | 150  | 0.7962         |
+| 4.0   | 160  | 0.7930         |
+| 5.0   | 200  | 0.8536         |
+| 6.0   | 240  | 0.8263         |
+| 6.25  | 250  | 0.8257         |
+| 7.0   | 280  | 0.8475         |
+| 7.5   | 300  | 0.8505         |
+| 8.0   | 320  | 0.8499         |
+| 8.75  | 350  | 0.8582         |
+| 9.0   | 360  | 0.8576         |
+| 10.0  | 400  | 0.8576         |
+### Framework Versions
+- Python: 3.11.11
+- Sentence Transformers: 3.4.1
+- Transformers: 4.48.2
+- PyTorch: 2.5.1+cu124
+- Accelerate: 1.3.0
+- Datasets: 3.2.0
+- Tokenizers: 0.21.0
+## Citation
+### BibTeX
+#### Sentence Transformers
+```bibtex
+@inproceedings{reimers-2019-sentence-bert,
+    title = "Sentence-BERT: Sentence Embeddings using Siamese BERT-Networks",
+    author = "Reimers, Nils and Gurevych, Iryna",
+    booktitle = "Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing",
+    month = "11",
+    year = "2019",
+    publisher = "Association for Computational Linguistics",
+    url = "https://arxiv.org/abs/1908.10084",
+}
+```
+#### MatryoshkaLoss
+```bibtex
+@misc{kusupati2024matryoshka,
+    title={Matryoshka Representation Learning},
+    author={Aditya Kusupati and Gantavya Bhatt and Aniket Rege and Matthew Wallingford and Aditya Sinha and Vivek Ramanujan and William Howard-Snyder and Kaifeng Chen and Sham Kakade and Prateek Jain and Ali Farhadi},
+    year={2024},
+    eprint={2205.13147},
+    archivePrefix={arXiv},
+    primaryClass={cs.LG}
+}
+```
+#### MultipleNegativesRankingLoss
+```bibtex
+@misc{henderson2017efficient,
+    title={Efficient Natural Language Response Suggestion for Smart Reply},
+    author={Matthew Henderson and Rami Al-Rfou and Brian Strope and Yun-hsuan Sung and Laszlo Lukacs and Ruiqi Guo and Sanjiv Kumar and Balint Miklos and Ray Kurzweil},
+    year={2017},
+    eprint={1705.00652},
+    archivePrefix={arXiv},
+    primaryClass={cs.CL}
+}
+```
+<!--
+## Glossary
+*Clearly define terms in order to be accessible across audiences.*
+-->
+<!--
+## Model Card Authors
+*Lists the people who create the model card, providing recognition and accountability for the detailed work that goes into its construction.*
+-->
+<!--
+## Model Card Contact
+*Provides a way for people who have updates to the Model Card, suggestions, or questions, to contact the Model Card authors.*
+-->

config.json ADDED Viewed

	@@ -0,0 +1,25 @@

+{
+  "_name_or_path": "Snowflake/snowflake-arctic-embed-l",
+  "architectures": [
+    "BertModel"
+  ],
+  "attention_probs_dropout_prob": 0.1,
+  "classifier_dropout": null,
+  "hidden_act": "gelu",
+  "hidden_dropout_prob": 0.1,
+  "hidden_size": 1024,
+  "initializer_range": 0.02,
+  "intermediate_size": 4096,
+  "layer_norm_eps": 1e-12,
+  "max_position_embeddings": 512,
+  "model_type": "bert",
+  "num_attention_heads": 16,
+  "num_hidden_layers": 24,
+  "pad_token_id": 0,
+  "position_embedding_type": "absolute",
+  "torch_dtype": "float32",
+  "transformers_version": "4.48.2",
+  "type_vocab_size": 2,
+  "use_cache": true,
+  "vocab_size": 30522
+}

config_sentence_transformers.json ADDED Viewed

	@@ -0,0 +1,12 @@

+{
+  "__version__": {
+    "sentence_transformers": "3.4.1",
+    "transformers": "4.48.2",
+    "pytorch": "2.5.1+cu124"
+  },
+  "prompts": {
+    "query": "Represent this sentence for searching relevant passages: "
+  },
+  "default_prompt_name": null,
+  "similarity_fn_name": "cosine"
+}

model.safetensors ADDED Viewed

	@@ -0,0 +1,3 @@

+version https://git-lfs.github.com/spec/v1
+oid sha256:007afa097a0e16fe9d16b2a3c72b85a55d3e55e96c83773daaa93d600731967e
+size 1336413848

modules.json ADDED Viewed

	@@ -0,0 +1,20 @@

+[
+  {
+    "idx": 0,
+    "name": "0",
+    "path": "",
+    "type": "sentence_transformers.models.Transformer"
+  },
+  {
+    "idx": 1,
+    "name": "1",
+    "path": "1_Pooling",
+    "type": "sentence_transformers.models.Pooling"
+  },
+  {
+    "idx": 2,
+    "name": "2",
+    "path": "2_Normalize",
+    "type": "sentence_transformers.models.Normalize"
+  }
+]

sentence_bert_config.json ADDED Viewed

	@@ -0,0 +1,4 @@

+{
+  "max_seq_length": 512,
+  "do_lower_case": false
+}

special_tokens_map.json ADDED Viewed

	@@ -0,0 +1,37 @@

+{
+  "cls_token": {
+    "content": "[CLS]",
+    "lstrip": false,
+    "normalized": false,
+    "rstrip": false,
+    "single_word": false
+  },
+  "mask_token": {
+    "content": "[MASK]",
+    "lstrip": false,
+    "normalized": false,
+    "rstrip": false,
+    "single_word": false
+  },
+  "pad_token": {
+    "content": "[PAD]",
+    "lstrip": false,
+    "normalized": false,
+    "rstrip": false,
+    "single_word": false
+  },
+  "sep_token": {
+    "content": "[SEP]",
+    "lstrip": false,
+    "normalized": false,
+    "rstrip": false,
+    "single_word": false
+  },
+  "unk_token": {
+    "content": "[UNK]",
+    "lstrip": false,
+    "normalized": false,
+    "rstrip": false,
+    "single_word": false
+  }
+}

tokenizer.json ADDED Viewed

The diff for this file is too large to render. See raw diff

tokenizer_config.json ADDED Viewed

	@@ -0,0 +1,63 @@

+{
+  "added_tokens_decoder": {
+    "0": {
+      "content": "[PAD]",
+      "lstrip": false,
+      "normalized": false,
+      "rstrip": false,
+      "single_word": false,
+      "special": true
+    },
+    "100": {
+      "content": "[UNK]",
+      "lstrip": false,
+      "normalized": false,
+      "rstrip": false,
+      "single_word": false,
+      "special": true
+    },
+    "101": {
+      "content": "[CLS]",
+      "lstrip": false,
+      "normalized": false,
+      "rstrip": false,
+      "single_word": false,
+      "special": true
+    },
+    "102": {
+      "content": "[SEP]",
+      "lstrip": false,
+      "normalized": false,
+      "rstrip": false,
+      "single_word": false,
+      "special": true
+    },
+    "103": {
+      "content": "[MASK]",
+      "lstrip": false,
+      "normalized": false,
+      "rstrip": false,
+      "single_word": false,
+      "special": true
+    }
+  },
+  "clean_up_tokenization_spaces": true,
+  "cls_token": "[CLS]",
+  "do_lower_case": true,
+  "extra_special_tokens": {},
+  "mask_token": "[MASK]",
+  "max_length": 512,
+  "model_max_length": 512,
+  "pad_to_multiple_of": null,
+  "pad_token": "[PAD]",
+  "pad_token_type_id": 0,
+  "padding_side": "right",
+  "sep_token": "[SEP]",
+  "stride": 0,
+  "strip_accents": null,
+  "tokenize_chinese_chars": true,
+  "tokenizer_class": "BertTokenizer",
+  "truncation_side": "right",
+  "truncation_strategy": "longest_first",
+  "unk_token": "[UNK]"
+}

vocab.txt ADDED Viewed

The diff for this file is too large to render. See raw diff