Multi-task learning shares information between related tasks, sometimes the number of parameters required. State-of-the-art results across natural language understanding tasks in the GLUE benchmark have used transfer from a single large task: unsupervised pre-training BERT, where a separate BERT model was fine-tuned for each task. We explore-task approaches that share a single BERT model with a small number of task-specific parameters. Using new adaptation modules, PALs or`projected attention layers', we match the performance of separately fine-tuned on the GLUE benchmark with roughly 7 times fewer parameters, and obtain-of-the-art results on the Recognizing Textual Entailment dataset.
No takes yet. Share an insight, caveat, or question.
Stickland et al. (2019) studied this question.