Skip to main navigation
Skip to search
Skip to main content
Sort by:
Keyphrases
Curriculum Learning
100%
Knowledge Distillation
100%
Bidirectional Encoder Representations from Transformers
100%
Curriculum Knowledge
100%
Multi-exit
100%
Training Samples
66%
Distillation
66%
Training Paradigm
66%
Intermediate Features
66%
Textual Entailment
66%
Early Exit
66%
Benchmark Dataset
33%
Sample-level
33%
Two-level
33%
Superior Performance
33%
Real-world Deployment
33%
Scholarly Attention
33%
Feature Level
33%
Generalization Ability
33%
Training Process
33%
Sample Feature
33%
Number of Parameters
33%
Performance Deterioration
33%
Vanilla
33%
Answer Selection
33%
Distillation Learning
33%
Performance Reduction
33%
Multi-exit Architecture
33%
Engineering
Experimental Result
100%
Level Feature
100%
Deep Layer
100%
Performance Deterioration
100%
Knowledge Distillation
100%
Computer Science
Bidirectional Encoder Representations From Transformers
100%
Knowledge Distillation
100%
Training Sample
66%
Experimental Result
33%
Superior Performance
33%
Training Process
33%