Build error
Build error
File size: 11,056 Bytes
8c90e7d |
1 2 3 4 5 6 7 8 9 10 11 12 13 14 15 16 17 18 19 20 21 22 23 24 25 26 27 28 29 30 31 32 33 34 35 36 37 38 39 40 41 42 43 44 45 46 47 48 49 50 51 52 53 54 55 56 57 58 59 60 61 62 63 64 65 66 67 68 69 70 71 72 73 74 75 76 77 78 79 80 81 82 83 84 85 86 87 88 89 90 91 92 93 94 95 96 97 98 99 100 101 102 103 104 105 106 107 108 109 110 111 112 113 114 115 116 117 118 119 120 121 122 123 124 125 126 127 128 129 130 131 132 133 134 135 136 137 138 139 140 141 142 143 144 145 146 147 148 149 150 151 152 153 154 155 156 157 158 159 160 161 162 163 164 165 166 167 168 169 170 171 172 173 174 175 176 177 |
# Pay Less Attention with Lightweight and Dynamic Convolutions (Wu et al., 2019)
This page contains pointers to pre-trained models as well as instructions on how to train new models for [our paper](
## Citation:
title = {Pay Less Attention with Lightweight and Dynamic Convolutions},
author = {Felix Wu and Angela Fan and Alexei Baevski and Yann Dauphin and Michael Auli},
booktitle = {International Conference on Learning Representations},
year = {2019},
url = {},
## Translation
### Pre-trained models
For some datasets we release models without GLUs which are faster at inference.
Model | Description | Dataset | Download
`` | LightConv (without GLUs) | [IWSLT14 German-English]( | model: <br> [download (.tar.gz)]( <br> IWSLT14 test: <br> [download (.tar.bz2)](
`` | DynamicConv (without GLUs) | [IWSLT14 German-English]( | model: <br> [download (.tar.gz)]( <br> IWSLT14 test: <br> [download (.tar.bz2)](
`lightconv.no_glu.wmt16.en-de` | LightConv (without GLUs) | [WMT16 English-German]( | model: <br> [download (.tar.gz)]( <br> newstest2014 (shared vocab): <br> [download (.tar.bz2)](
`dynamicconv.no_glu.wmt16.en-de` | DynamicConv (without GLUs) | [WMT16 English-German]( | model: <br> [download (.tar.gz)]( <br> newstest2014 (shared vocab): <br> [download (.tar.bz2)](
`lightconv.glu.wmt16.en-de` | LightConv | [WMT16 English-German]( | model: <br> [download (.tar.gz)]( <br> newstest2014 (shared vocab): <br> [download (.tar.bz2)](
`dynamicconv.glu.wmt16.en-de` | DynamicConv | [WMT16 English-German]( | model: <br> [download (.tar.gz)]( <br> newstest2014 (shared vocab): <br> [download (.tar.bz2)](
`lightconv.glu.wmt14.en-fr` | LightConv | [WMT14 English-French]( | model: <br> [download (.tar.gz)]( <br> newstest2014: <br> [download (.tar.bz2)](
`dynamicconv.glu.wmt14.en-fr` | DynamicConv | [WMT14 English-French]( | model: <br> [download (.tar.gz)]( <br> newstest2014: <br> [download (.tar.bz2)](
`lightconv.glu.wmt17.zh-en` | LightConv | [WMT17 Chinese-English]( | model: <br> [download (.tar.gz)]( <br> newstest2017: <br> [download (.tar.bz2)](
`dynamicconv.glu.wmt17.zh-en` | DynamicConv | [WMT17 Chinese-English]( | model: <br> [download (.tar.gz)]( <br> newstest2017: <br> [download (.tar.bz2)](
### Memory-Efficient CUDA Kernels
Since the PyTorch implementations of Light/Dynamic conv are quite memory intensive, we have developed CUDA kernels that implement the light and dynamic convolution operator in a memory-efficient and performant manner. For large sequence lengths, these kernels save about 50% memory compared to the PyTorch equivalent.
To install the kernels, use the commands below. Once installed, they will automatically be used in place of the PyTorch implementations whenever a light or dynamic convolution is used.
# to install lightconv
cd fairseq/modules/lightconv_layer
python install
# to install dynamicconv
cd fairseq/modules/dynamicconv_layer
python install
### Example usage (torch.hub)
We require a few additional Python dependencies for preprocessing:
pip install sacremoses subword_nmt
Interactive translation via PyTorch Hub:
import torch
# List available models
torch.hub.list('pytorch/fairseq') # [..., 'lightconv.glu.wmt17.zh-en', ... ]
# Load a transformer trained on WMT'16 En-De
zh2en = torch.hub.load('pytorch/fairseq', 'lightconv.glu.wmt17.zh-en', tokenizer='moses', bpe='subword_nmt')
# The underlying model is available under the *models* attribute
assert isinstance(zh2en.models[0], fairseq.models.lightconv.LightConvModel)
# Translate a sentence
zh2en.translate('你好 世界')
# 'Hello World'
Loading custom models:
from fairseq.models.lightconv import LightConvModel
en2fr = LightConvModel.from_pretrained(
en2fr.translate('Hello world!')
# 'Bonjour le monde'
### Preprocessing the training datasets
Please follow the instructions in [`examples/translation/`](../translation/ to preprocess the data.
### Training and evaluation options:
To use the model without GLU, please set `--encoder-glu 0 --decoder-glu 0`.
For LightConv, please use `--encoder-conv-type lightweight --decoder-conv-type lightweight`, otherwise the default is DynamicConv.
For best BLEU results, lenpen may need to be manually tuned.
To use the CUDA kernels, first install the PyTorch modules using the commands
above. Once the CUDA modules are installed, they will automatically be used
instead of the PyTorch modules.
### IWSLT14 De-En
Training and evaluating DynamicConv (without GLU) on a GPU:
# Training
mkdir -p $SAVE
CUDA_VISIBLE_DEVICES=0 $(which fairseq-train) data-bin/ \
--clip-norm 0 --optimizer adam --lr 0.0005 \
--source-lang de --target-lang en --max-tokens 4000 --no-progress-bar \
--log-interval 100 --stop-min-lr '1e-09' --weight-decay 0.0001 \
--criterion label_smoothed_cross_entropy --label-smoothing 0.1 \
--lr-scheduler inverse_sqrt \
--ddp-backend=legacy_ddp \
--max-update 50000 --warmup-updates 4000 --warmup-init-lr '1e-07' \
--adam-betas '(0.9, 0.98)' --keep-last-epochs 10 \
-a lightconv_iwslt_de_en --save-dir $SAVE \
--dropout 0.3 --attention-dropout 0.1 --weight-dropout 0.1 \
--encoder-glu 0 --decoder-glu 0
python scripts/ --inputs $SAVE \
--num-epoch-checkpoints 10 --output "${SAVE}/"
# Evaluation
CUDA_VISIBLE_DEVICES=0 fairseq-generate data-bin/ --path "${SAVE}/" --batch-size 128 --beam 4 --remove-bpe --lenpen 1 --gen-subset test --quiet
### WMT16 En-De
Training and evaluating DynamicConv (with GLU) on WMT16 En-De using cosine scheduler on one machine with 8 V100 GPUs:
# Training
mkdir -p $SAVE
python -m torch.distributed.launch --nproc_per_node 8 $(which fairseq-train) \
data-bin/wmt16_en_de_bpe32k --fp16 --log-interval 100 --no-progress-bar \
--max-update 30000 --share-all-embeddings --optimizer adam \
--adam-betas '(0.9, 0.98)' --clip-norm 0.0 --weight-decay 0.0 \
--criterion label_smoothed_cross_entropy --label-smoothing 0.1 \
--stop-min-lr 1e-09 --update-freq 16 --attention-dropout 0.1 --keep-last-epochs 10 \
--ddp-backend=legacy_ddp --max-tokens 3584 \
--lr-scheduler cosine --warmup-init-lr 1e-7 --warmup-updates 10000 \
--lr-shrink 1 --lr 0.001 --min-lr 1e-7 --warmup-init-lr 1e-07 \
--t-mult 1 --lr-period-updates 20000 \
--arch lightconv_wmt_en_de_big --save-dir $SAVE \
--dropout 0.3 --attention-dropout 0.1 --weight-dropout 0.1 \
--encoder-glu 1 --decoder-glu 1
# Evaluation
CUDA_VISIBLE_DEVICES=0 fairseq-generate data-bin/wmt16.en-de.joined-dict.newstest2014 --path "${SAVE}/" --batch-size 128 --beam 5 --remove-bpe --lenpen 0.5 --gen-subset test > wmt16_gen.txt
bash scripts/ wmt16_gen.txt
### WMT14 En-Fr
Training DynamicConv (with GLU) on WMT14 En-Fr using cosine scheduler on one machine with 8 V100 GPUs:
# Training
mkdir -p $SAVE
python -m torch.distributed.launch --nproc_per_node 8 $(which fairseq-train) \
data-bin/wmt14_en_fr --fp16 --log-interval 100 --no-progress-bar \
--max-update 30000 --share-all-embeddings --optimizer adam \
--adam-betas '(0.9, 0.98)' --clip-norm 0.0 --weight-decay 0.0 \
--criterion label_smoothed_cross_entropy --label-smoothing 0.1 \
--stop-min-lr 1e-09 --update-freq 16 --attention-dropout 0.1 --keep-last-epochs 10 \
--ddp-backend=legacy_ddp --max-tokens 3584 \
--lr-scheduler cosine --warmup-init-lr 1e-7 --warmup-updates 10000 \
--lr-shrink 1 --lr 0.001 --min-lr 1e-7 --warmup-init-lr 1e-07 \
--t-mult 1 --lr-period-updates 70000 \
--arch lightconv_wmt_en_fr_big --save-dir $SAVE \
--dropout 0.1 --attention-dropout 0.1 --weight-dropout 0.1 \
--encoder-glu 1 --decoder-glu 1
# Evaluation
CUDA_VISIBLE_DEVICES=0 fairseq-generate data-bin/wmt14.en-fr.joined-dict.newstest2014 --path "${SAVE}/" --batch-size 128 --beam 5 --remove-bpe --lenpen 0.9 --gen-subset test