Non-autoregressive Translation with Dependency-Aware Decoder

Jiaao Zhan,Qian Chen,Boxing Chen,Wen Wang,Yu Bai,Yang Gao
DOI: https://doi.org/10.48550/arxiv.2203.16266
2022-01-01
Abstract: Non-autoregressive translation (NAT) models suffer from inferior translation quality due to removal of dependency on previous target tokens from inputs to the decoder. In this paper, we propose a novel and general approach to enhance the target dependency within the NAT decoder from two perspectives: decoder input and decoder self-attention. First, we transform the initial decoder input from the source language space to the target language space through a novel attentive transformation process. The transformation reassembles the decoder input based on target token embeddings and conditions the final output on the target-side information. Second, before NAT training, we introduce an effective forward-backward pre-training phase, implemented with different triangle attention masks. This pre-training phase enables the model to gradually learn bidirectional dependencies for the final NAT decoding process. Experimental results demonstrate that the proposed approaches consistently improve highly competitive NAT models on four WMT translation directions by up to 1.88 BLEU score, while overall maintaining inference latency comparable to other fully NAT models.
What problem does this paper attempt to address?