Abstract:We have witnessed superhuman intelligence thanks to the fast development of large language models and multimodal language models. As the application of such superhuman models becomes more and more common, a critical question rises here: how can we ensure superhuman models are still safe, reliable and aligned well to human values? In this position paper, we discuss the concept of superalignment from the learning perspective to answer this question by outlining the learning paradigm shift from large-scale pretraining, supervised fine-tuning, to alignment training. We define superalignment as designing effective and efficient alignment algorithms to learn from noisy-labeled data (point-wise samples or pair-wise preference data) in a scalable way when the task becomes very complex for human experts to annotate and the model is stronger than human experts. We highlight some key research problems in superalignment, namely, weak-to-strong generalization, scalable oversight, and evaluation. We then present a conceptual framework for superalignment, which consists of three modules: an attacker which generates adversary queries trying to expose the weaknesses of a learner model; a learner which will refine itself by learning from scalable feedbacks generated by a critic model along with minimal human experts; and a critic which generates critics or explanations for a given query-response pair, with a target of improving the learner by criticizing. We discuss some important research problems in each component of this framework and highlight some interesting research ideas that are closely related to our proposed framework, for instance, self-alignment, self-play, self-refinement, and more. Last, we highlight some future research directions for superalignment, including identification of new emergent risks and multi-dimensional alignment.

Improving Weak-to-Strong Generalization with Reliability-Aware Alignment

Weak-to-Strong Generalization beyond Accuracy: a Pilot Study in Safety, Toxicity, and Legal Reasoning

Super(ficial)-alignment: Strong Models May Deceive Weak Models in Weak-to-Strong Generalization

A transfer learning framework for weak-to-strong generalization

Improving Weak-to-Strong Generalization with Scalable Oversight and Ensemble Learning

Improving the Robustness of Large Language Models via Consistency Alignment

Improving Generalization of Alignment with Human Preferences through Group Invariant Learning

Weak-to-Strong Generalization: Eliciting Strong Capabilities With Weak Supervision

On the Calibration of Large Language Models and Alignment

Your Weak LLM is Secretly a Strong Teacher for Alignment

Vision Superalignment: Weak-to-Strong Generalization for Vision Foundation Models

The Superalignment of Superhuman Intelligence with Large Language Models

How Reliable Is Human Feedback For Aligning Large Language Models?

Human-Instruction-Free LLM Self-Alignment with Limited Samples

Co-Supervised Learning: Improving Weak-to-Strong Generalization with Hierarchical Mixture of Experts

Explanation, Debate, Align: A Weak-to-Strong Framework for Language Model Generalization

Not Everything is All You Need: Toward Low-Redundant Optimization for Large Language Model Alignment

Trustworthy LLMs: a Survey and Guideline for Evaluating Large Language Models' Alignment

Decoupled Alignment for Robust Plug-and-Play Adaptation

Weak Alignment Supervision from Hybrid Model Improves End-to-end ASR

Weak-to-Strong Preference Optimization: Stealing Reward from Weak Aligned Model