Abstract:Machine learning (ML) has been increasingly used in a variety of domains, while solving ML programming tasks poses unique challenges due to the fundamental difference in the nature and the construct of general programming tasks, especially for developers who do not have ML backgrounds. Automatic code generation that produces a code snippet from a natural language description can be a promising technique to accelerate ML programming tasks. In recent years, although many deep learning-based neural code generation models have been proposed with high accuracy, the fact that most of them are mainly evaluated on general programming tasks calls into question their effectiveness and usefulness in ML programming tasks. In this paper, we set out to investigate the effectiveness of existing neural code generation models on ML programming tasks. For our analysis, we select six state-of-the-art neural code generation models, and evaluate their performance on four widely used ML libraries, with newly created 83K pairs of natural-language described ML programming tasks. Our empirical study reveals some good, bad, and missing aspects of neural code generation models on ML tasks, with a few major ones listed below. ( Good ) Neural code generation models perform significantly better on ML tasks than on non-ML tasks with an average difference of 10.6 points in BLEU-4 scores. ( Bad ) More than 80% of the generated code is semantically incorrect. ( Bad ) Code generation models do not have significance in improving developers’ completion time. ( Good ) The generated code can help developers write correct code by providing developers with clues for using correct APIs. ( Missing ) The observation from our user study reveals the missing aspects of code generation for ML tasks, e.g., decomposing code generation for divide-and-conquer into API sequence identification and API usage generation.

Seq2Seq or Seq2Tree: Generating Code Using Both Paradigms Via Mutual Learning.

From Code to Natural Language: Type-Aware Sketch-Based Seq2Seq Learning

CODEP: Grammatical Seq2Seq Model for General-Purpose Code Generation.

code2seq: Generating Sequences from Structured Representations of Code

Bridging Code Semantic and LLMs: Semantic Chain-of-Thought Prompting for Code Generation

CodeTree: Agent-guided Tree Search for Code Generation with Large Language Models

Leveraging Code Generation to Improve Code Retrieval and Summarization via Dual Learning

Code Generation with Hybrid of Structural and Semantic Features Retrieval

A syntax-guided multi-task learning approach for Turducken-style code generation

Code Generation as a Dual Task of Code Summarization

Antecedent Predictions Are More Important Than You Think: An Effective Method for Tree-Based Code Generation

CodeGRAG: Bridging the Gap between Natural Language and Programming Language via Graphical Retrieval Augmented Generation

TreeGen: A Tree-Based Transformer Architecture for Code Generation

GAP-Gen: Guided Automatic Python Code Generation

Automatic Code Annotation Generation Based on Multi-dimensional Heterogeneous Graph Structure

Neural Comment Generation for Source Code with Auxiliary Code Classification Task

Tree-Structured Neural Machine for Linguistics-Aware Sentence Generation

Automated Source Code Generation and Auto-completion Using Deep Learning: Comparing and Discussing Current Language-Model-Related Approaches

The Good, the Bad, and the Missing: Neural Code Generation for Machine Learning Tasks

Think Outside the Code: Brainstorming Boosts Large Language Models in Code Generation

Multi-Programming Language Ensemble for Code Generation in Large Language Model