Abstract:LLMs have revolutionized the field of artificial intelligence and have emerged as the de-facto tool for many tasks. The current established technology of LLMs is to process input and generate output at the token level. This is in sharp contrast to humans who operate at multiple levels of abstraction, well beyond single words, to analyze information and to generate creative content. In this paper, we present an attempt at an architecture which operates on an explicit higher-level semantic representation, which we name a concept. Concepts are language- and modality-agnostic and represent a higher level idea or action in a flow. Hence, we build a "Large Concept Model". In this study, as proof of feasibility, we assume that a concept corresponds to a sentence, and use an existing sentence embedding space, SONAR, which supports up to 200 languages in both text and speech modalities. The Large Concept Model is trained to perform autoregressive sentence prediction in an embedding space. We explore multiple approaches, namely MSE regression, variants of diffusion-based generation, and models operating in a quantized SONAR space. These explorations are performed using 1.6B parameter models and training data in the order of 1.3T tokens. We then scale one architecture to a model size of 7B parameters and training data of about 2.7T tokens. We perform an experimental evaluation on several generative tasks, namely summarization and a new task of summary expansion. Finally, we show that our model exhibits impressive zero-shot generalization performance to many languages, outperforming existing LLMs of the same size. The training code of our models is freely available.

On the Tip of the Tongue: Analyzing Conceptual Representation in Large Language Models with Reverse-Dictionary Probe

Concise and Organized Perception Facilitates Large Language Models for Deductive Reasoning.

Conceptual and Unbiased Reasoning in Language Models

Probing Conceptual Understanding of Large Visual-Language Models

Large Language Models Are In-Context Semantic Reasoners Rather Than Symbolic Reasoners

A prompt construction method for the reverse dictionary task of large-scale language models

Exploring Concept Depth: How Large Language Models Acquire Knowledge at Different Layers?

Towards Concept-Aware Large Language Models

COPEN: Probing Conceptual Knowledge in Pre-trained Language Models

Probing Linguistic Information For Logical Inference In Pre-trained Language Models

A Latent-Variable Model for Intrinsic Probing

Interventional Probing in High Dimensions: An NLI Case Study

How Large Language Models Encode Context Knowledge? A Layer-Wise Probing Study

Inductive Linguistic Reasoning with Large Language Models

Thought Propagation: An Analogical Approach to Complex Reasoning with Large Language Models

Exploring Multilingual Probing in Large Language Models: A Cross-Language Analysis

Chain-of-Thought in Large Language Models: Decoding, Projection, and Activation

Large Concept Models: Language Modeling in a Sentence Representation Space

Large Language Models as Analogical Reasoners

Probing Language Models on Their Knowledge Source