Artificial intelligence (AI)
Field of research and engineering that aims to build systems capable of performing tasks such as perceiving, predicting, generating or deciding from inputs.
PANACHES Knowledge
This glossary explains the essential concepts of artificial intelligence, from machine learning and generative models to data, agents, evaluation and safety. It is intended for curious readers, creat…
595 terms to explore
595 terms
24 terms
Field of research and engineering that aims to build systems capable of performing tasks such as perceiving, predicting, generating or deciding from inputs.
Functional assembly combining a model, software, data, interfaces and possibly actions to produce predictions, recommendations, content or decisions.
Finite sequence of defined instructions that transforms inputs into outputs; not all algorithms fall within AI.
Mathematical or computational representation whose parameters and rules process inputs and produce outputs.
Approach based on explicit rules, logic and knowledge representation, as opposed to statistical learning.
Approach based on artificial neural networks whose parameters are learned from data.
Approach relying on probabilistic models and statistical inference to learn from data and generalize.
AI systems that produce new content (text, images, audio, code, video) rather than only classifying or predicting labels.
AI systems that estimate a likely outcome from inputs, often via supervised learning.
Models that learn the boundary between classes or predict conditional outputs, as opposed to modeling the full data distribution.
Hypothetical AI matching the breadth and adaptability of human cognition across most tasks; not yet realized.
Hypothetical AI whose capabilities would exceed the best human performance in virtually all domains.
AI designed to perform a specific task or a limited set of tasks, without general cognitive abilities.
Use of systems to execute tasks with reduced human intervention; AI is one means among others to automate.
Ability of a system to plan and execute actions toward a goal with limited human oversight; AI autonomy is graded, not binary.
Software that encodes domain knowledge as rules and uses an inference engine to derive conclusions or recommendations.
Structured set of facts and rules used by a symbolic AI system to reason and produce answers.
Component of an expert system that applies logical rules to a knowledge base to derive conclusions.
Exploration of a state space guided by heuristic functions to find a solution efficiently without exhaustive enumeration.
Rule of thumb or approximation that guides search or decision-making without guaranteeing an optimal result.
Set of all configurations reachable from an initial state by applying the available actions; used to formalize search and planning problems.
Function that maps a solution to a score to be optimized (minimized or maximized); also called loss or cost in training.
In AI, the phase where a trained model produces outputs from new inputs; in logic, the derivation of conclusions from premises.
Ability that appears in a model only above a certain scale or complexity threshold and was not explicitly programmed.
14 terms
Proposal by Alan Turing to assess whether a machine can exhibit behavior indistinguishable from a human in conversation.
1956 summer workshop at Dartmouth College that gave birth to the field of artificial intelligence as an academic discipline.
Period of reduced funding and interest in AI following unmet expectations, occurring in the late 1980s and 1990s.
Good Old-Fashioned AI: symbolic, rule-based AI approaches dominant before the rise of statistical and neural methods.
Single-layer linear classifier introduced by Frank Rosenblatt in 1958, a precursor of modern neural networks.
Period beginning in the 2010s when deep neural networks, big data and GPU compute re-established connectionist AI as the dominant paradigm.
Machine learning using neural networks with many layers, capable of learning hierarchical representations from large datasets.
Large model trained on broad data at scale and adaptable to many downstream tasks via fine-tuning or prompting.
AI model that displays significant generality and can perform a wide range of tasks; term used in the EU AI Act.
Current state of the art in model capability, often pushed by the largest foundation models.
Approach where AI assists humans in decision-making rather than replacing them, emphasizing human-AI collaboration.
System combining symbolic reasoning and neural or statistical components to leverage strengths of both paradigms.
Approach integrating neural networks with symbolic representations or logic to improve reasoning and interpretability.
AI models running directly on embedded devices (microcontrollers, edge hardware) rather than in the cloud.
25 terms
Single numerical value, as opposed to a vector or tensor.
Ordered array of numbers representing a point or direction in a multidimensional space; central object in ML computations.
Two-dimensional array of numbers; basic structure for linear transformations and batch computations in ML.
Multidimensional array generalizing scalars, vectors and matrices; the core data structure of deep learning frameworks.
Number of components of a tensor, or size of a vector space; high-dimensional spaces underlie embeddings and latent representations.
Lower-dimensional representation space learned by a model in which similar inputs are mapped to nearby points.
Variable whose value depends on the outcome of a random phenomenon; central to probabilistic modeling.
Function describing the probabilities of the possible values of a random variable.
Probability of an event given that another event has occurred; written P(A|B).
Probability distribution expressing beliefs about a quantity before observing data; combined with the likelihood to form the posterior.
Probability distribution of a quantity after combining the prior with the observed data via Bayes' rule.
Method that updates beliefs about unknown quantities using Bayes' rule, combining prior and likelihood.
Probability of the observed data given a set of parameters; used in Bayesian inference and maximum likelihood estimation.
Function measuring the discrepancy between model predictions and targets; minimized during training.
Vector of partial derivatives indicating the direction of steepest increase of a function; used in reverse for gradient descent.
Derivative of a multivariate function with respect to one variable, treating the others as constant; component of the gradient.
Optimization algorithm that updates parameters in the direction opposite to the gradient of the loss to minimize it.
Optimization problem whose objective and constraints define a convex set, guaranteeing that a local optimum is global.
Point where the loss is lower than in a neighborhood, without being the global minimum; optimization can get trapped there.
Technique that penalizes model complexity (e.g., L2, dropout) to improve generalization and prevent overfitting.
Function measuring the size of a vector (e.g., L1, L2); used in regularization and similarity computations.
Similarity measure between two vectors based on the cosine of the angle between them; widely used for embeddings.
Standard geometric distance between two points in a vector space, computed as the L2 norm of their difference.
Measure of uncertainty or information content of a probability distribution.
Loss function measuring the difference between two probability distributions; standard loss for classification.
33 terms
Examples used to fit model parameters during training.
Subset of the data used to train the model.
Subset of the data used to tune hyperparameters and monitor generalization during training, without touching the test set.
Subset of the data held out to evaluate the final model's generalization, used only once.
Structured collection of examples used to train, validate or test a model.
Large structured collection of texts or other data used for training or linguistic analysis.
Single instance in a dataset, typically a pair (input, target) or (input, label).
Individual measurable variable used as input by a model.
Variable to predict, also called label in classification.
Discrete or continuous value assigned to an example as the prediction target, especially in classification.
Label or metadata added to a data sample, usually by human annotators, to create a supervised dataset.
Verified reference value or label considered correct for evaluation and supervised training.
Description of the structure of a dataset: fields, types, relationships.
Structured information describing a dataset or model: source, license, creation date, intended use, etc.
Data organized in a defined schema such as tables, JSON or XML.
Data without a predefined schema, such as raw text, images, audio or video.
Data with partial structure, such as JSON, XML or logs, that does not fit a strict relational schema.
Transformations applied to raw data before training or inference (cleaning, normalization, tokenization, etc.).
Detection and correction of errors, inconsistencies and duplicates in a dataset.
Removal of duplicate examples from a dataset; critical for training corpora, especially for LLMs.
Rescaling of features to a common range (e.g., [0, 1]) to stabilize training.
Centering and scaling features to zero mean and unit variance.
Situation where some classes are much more frequent than others in a classification dataset; biases models toward the majority class.
Selection of a subset from a population or dataset according to a procedure (random, stratified, etc.).
Sampling method that preserves the proportion of each class or subgroup in the selected subset.
Generation of new training examples by applying label-preserving transformations (crop, noise, paraphrase, etc.) to existing data.
Artificially generated data, often by a model, used to augment or replace real data.
Data requiring special protection, such as personal data, health data or confidential business information.
Unintended leakage of information from training into validation or test, producing overly optimistic metrics.
Presence of test data in the training corpus, which inflates evaluation metrics and invalidates them.
Change over time in the statistical distribution of input data relative to the training distribution.
Change over time in the relationship between inputs and the target concept to predict.
Traceability of the origin, transformations and uses of a dataset throughout its lifecycle.
30 terms
Set of methods that enable a system to learn patterns from data rather than from explicitly programmed rules.
Learning from labeled examples, where each training input is paired with a target output.
Learning structure from unlabeled data, for clustering, dimensionality reduction or density estimation.
Learning that combines a small amount of labeled data with a large amount of unlabeled data.
Learning where supervision signals are derived from the data itself, e.g., by predicting masked parts; foundation of modern foundation models.
Learning strategy where the model queries an oracle (often a human) to label the most informative examples, reducing annotation cost.
Learning where an agent learns a policy by interacting with an environment and receiving rewards or penalties.
Reuse of a model trained on one task or domain as a starting point for another task, to reduce data and compute needs.
Learning where the model is updated continuously as new data arrives, rather than in offline batches.
Sequential learning from a stream of tasks without forgetting previously acquired capabilities; addresses catastrophic forgetting.
Learning where multiple clients train locally and only share model updates, aggregated by a central server, preserving data privacy.
Supervised task of assigning a category to an input among a finite set of classes.
Supervised task of predicting a continuous numerical value from inputs.
Unsupervised task of grouping similar examples together without predefined labels.
Techniques that project high-dimensional data into a lower-dimensional space while preserving structure (e.g., PCA, UMAP, t-SNE).
Linear dimensionality reduction that projects data onto orthogonal axes of maximum variance.
Non-parametric method that classifies or regresses an input by the majority or average of its k closest training examples.
Model that recursively splits the input space according to feature thresholds to predict a target.
Ensemble of decision trees trained on bootstrap samples with random feature subsets; reduces variance and overfitting.
Ensemble method that sequentially trains weak models, each correcting the errors of the previous ones.
Boosting variant that fits each new model to the gradient of the loss with respect to the current predictions (e.g., XGBoost, LightGBM).
Classifier that finds the hyperplane maximizing the margin between classes in feature space.
Linear classifier modeling the log-odds of a class as a linear combination of features; baseline for binary classification.
Probabilistic classifier assuming conditional independence of features given the class; simple and effective on text.
Model combining the predictions of several base models to improve accuracy and robustness.
Situation where a model fits the training data too closely, including its noise, and generalizes poorly.
Situation where a model is too simple to capture the underlying structure of the data, performing poorly even on training.
Ability of a model to produce accurate predictions on new data drawn from the same distribution as the training data.
Fundamental tradeoff between the bias of a model (systematic error) and its variance (sensitivity to the training sample).
Evaluation method that splits the data into k folds and trains on k-1 to test on the remaining one, rotating to obtain all scores.
32 terms
Computing model composed of layers of interconnected artificial neurons, trained by gradient descent.
Basic unit of a neural network that computes a weighted sum of its inputs followed by a nonlinear activation function.
Set of neurons operating at the same level of abstraction in a neural network.
Trainable parameters of a neural network that scale the inputs of a neuron; learned by optimization.
Trainable additive parameter of a neuron, independent of its inputs.
Nonlinear function applied to the weighted sum in a neuron (ReLU, GELU, sigmoid, etc.).
Rectified Linear Unit: activation function returning max(0, x); standard in modern deep networks.
Gaussian Error Linear Unit: smooth activation used in Transformers, weighting inputs by their Gaussian CDF.
S-shaped activation function squashing its input to (0, 1); used for binary outputs and gating.
Function converting a vector of real values into a probability distribution summing to 1; used at the output of classifiers.
Phase of neural network computation that propagates inputs through the layers to produce outputs.
Algorithm that computes gradients of the loss with respect to each parameter by chain rule, enabling gradient descent.
Neural network using convolution layers, particularly effective for images, audio and other signals with local structure.
Operation applying a kernel sliding over the input to extract local features; core of CNNs.
Downsampling operation (max or average) that reduces spatial dimensions and adds some translation invariance.
Neural network processing sequences by maintaining a hidden state updated at each step; largely replaced by Transformers.
Long Short-Term Memory: RNN variant with gating mechanisms that mitigate the vanishing gradient problem.
Gated Recurrent Unit: simplified RNN variant combining the forget and input gates into a single update gate.
Neural network trained to reconstruct its input through a bottleneck, learning a compressed latent representation.
Generative autoencoder learning a probabilistic latent space enabling sampling and interpolation.
Generative model trained by adversarial competition between a generator and a discriminator.
Neural network operating on graph-structured data by aggregating neighbor information.
Deep architecture using skip connections to ease gradient flow and enable very deep networks.
Shortcut adding the input of a layer to its output, enabling residual learning and stabilizing deep training.
Normalization across the features of a single example; standard in Transformers.
Normalization across the examples of a mini-batch; stabilizes and accelerates training of deep networks.
Regularization technique that randomly drops units during training to prevent co-adaptation and overfitting.
Model that processes and relates several modalities (text, image, audio, etc.) in a shared representation space.
Architecture that routes each input to a subset of specialized expert sub-networks, scaling capacity without proportional compute cost.
Mechanism selecting which expert sub-networks process each input in an MoE model.
Training of a smaller student model to mimic a larger teacher model, often improving efficiency.
31 terms
Neural architecture based on self-attention, processing sequences in parallel and underlying most modern LLMs.
Mechanism weighting the importance of different positions or inputs in producing an output; central to Transformers.
Attention where queries, keys and values come from the same sequence, modeling internal dependencies.
Attention where queries come from one sequence and keys/values from another; used in encoder-decoder models and multimodal fusion.
Attention computed in parallel over several subspaces, allowing the model to attend to different relationships.
One of the parallel attention subspaces in multi-head attention.
Component of a model that builds a representation of the input; in Transformers, uses bidirectional self-attention.
Component of a model that generates outputs from a representation, typically autoregressively.
Architecture with separate encoder and decoder, classic for machine translation and many seq2seq tasks.
Architecture using only a decoder with causal attention; standard for modern LLMs such as GPT.
Atomic unit processed by a language model, typically a subword or word piece obtained by tokenization.
Process of splitting text into tokens according to a model's vocabulary.
Component that converts text into a sequence of tokens and back; defines the model's vocabulary.
Token unit smaller than a word, balancing vocabulary size and coverage of rare words (BPE, SentencePiece, WordPiece).
Subword tokenization algorithm that iteratively merges the most frequent pairs of symbols to build a vocabulary.
Reserved token with a specific role in a model: BOS, EOS, PAD, MASK, system, etc.
Learned dense vector representation of a token, word, sentence or other entity in a continuous space.
Vector representation of a token in the model's input space; learned during training.
Vector representation of an entire sentence or paragraph, used for semantic similarity or retrieval.
Signal added to token embeddings to inject position information into the model, since attention is permutation-equivariant.
Positional encoding that rotates query and key vectors by position-dependent angles; widely used in modern LLMs.
Matrix preventing attention to certain positions (padding, future tokens, etc.).
Attention mask preventing a token from attending to future tokens; essential for autoregressive language models.
Maximum number of tokens a model can process in a single forward pass, including prompt and generated output.
Capability to process very long inputs (hundreds of thousands of tokens) within the context window.
Caching of previously computed key and value tensors during autoregressive generation to avoid redundant computation.
Raw output of a model before the softmax, representing unnormalized scores for each possible token.
Probability distribution over the vocabulary produced by a language model for the next token.
Core training objective of autoregressive language models: predicting the next token given previous ones.
Training objective where some tokens in the input are masked and the model learns to predict them; used by BERT-style encoders.
Generation that produces outputs sequentially, each conditioned on previously generated outputs.
35 terms
Process of adjusting a model's parameters on a dataset to minimize a loss function.
Initial training of a model on a large corpus to learn general capabilities before fine-tuning for specific tasks.
Additional training of a pretrained model on a smaller, task-specific dataset to specialize it.
Fine-tuning of a pretrained model on labeled instruction-response pairs to make it follow instructions.
Model parameter updated by gradient descent during training, as opposed to frozen parameters.
Configuration of the training process (learning rate, batch size, number of layers) not learned by gradient descent.
Hyperparameter scaling the magnitude of parameter updates at each optimization step.
Strategy adjusting the learning rate during training (warmup, cosine decay, etc.) to improve convergence.
One complete pass of the training set through the optimization algorithm.
Single update of model parameters from one mini-batch.
Set of examples processed together in one forward/backward pass before a parameter update.
Small batch used in stochastic gradient descent, balancing update noise and computational efficiency.
Number of examples in a mini-batch; key hyperparameter affecting stability and throughput.
Technique accumulating gradients over several forward/backward passes before applying an update, simulating larger batch sizes.
Algorithm that updates model parameters from gradients (SGD, Adam, AdamW, etc.).
Stochastic Gradient Descent: optimizer that updates parameters using gradients computed on mini-batches, optionally with momentum.
Adaptive optimizer combining momentum and per-parameter learning rates; standard for deep learning.
Adam variant decoupling weight decay from the gradient update, improving generalization; standard for Transformers.
Regularization penalty proportional to parameter magnitude, encouraging smaller weights and better generalization.
Capping gradient norm or value to prevent exploding gradients in deep networks.
Phase at the start of training during which the learning rate is gradually increased from zero to its target value.
Snapshot of model weights (and optionally optimizer state) saved during training, enabling resumption.
Stopping of training when validation performance stops improving, to prevent overfitting.
Training spread across multiple GPUs or machines to handle models and datasets too large for a single device.
Each worker holds a full copy of the model and trains on a shard of the data, synchronizing gradients.
Model is split across multiple devices because it does not fit in memory on a single one.
Training using FP16/BF16 and FP32 together to reduce memory and accelerate computation while preserving stability.
Loss measured on the validation set; tracked to detect overfitting and tune hyperparameters.
Tendency of neural networks to abruptly lose previously learned capabilities when trained on new tasks.
Training a model from scratch on a new dataset, as opposed to fine-tuning a pretrained model.
Family of methods that adapt large models by training only a small number of additional parameters (LoRA, adapters, prompts).
Low-Rank Adaptation: PEFT method that injects trainable low-rank matrices into model weights, drastically reducing the number of trainable parameters.
LoRA combined with 4-bit quantization of the base model, enabling fine-tuning of large models on a single GPU.
Small trainable module inserted into a frozen pretrained model to specialize it for a new task.
Combining several trained adapters (e.g., for different tasks or languages) into a single set of weights.
14 terms
Process of steering a model's behavior to match human intentions, values and safety constraints.
Fine-tuning a model on a curated set of (instruction, response) pairs so it better follows natural-language instructions.
Pairwise or graded judgments of human annotators comparing model outputs, used to train reward models.
Model trained on human preferences to predict which of two outputs a human would prefer; central to RLHF.
Reinforcement Learning from Human Feedback: alignment technique that fine-tunes a model with RL against a reward model trained on human preferences.
Reinforcement Learning from AI Feedback: variant of RLHF where the reward signal is produced by another AI system.
Alignment method that directly optimizes a model on human preferences without training an explicit reward model.
Proximal Policy Optimization: RL algorithm widely used in RLHF for its stability.
Technique that samples multiple outputs from a model and keeps only the best ones (e.g., by reward) for fine-tuning.
Alignment approach where an AI critiques and revises its own outputs against a set of written principles.
Distillation of a stronger model's chain-of-thought reasoning into a smaller student model.
Phenomenon where excessive optimization against a proxy reward degrades true quality, a classic Goodhart's law effect.
Exploitation of flaws in a reward signal by an optimizing model, leading to behaviors that score high but are not desired.
Ability of a model to decline requests it should not fulfill while still answering legitimate ones.
30 terms
Model that assigns probabilities to sequences of tokens, often used to generate text.
Field of AI focused on understanding, generating and manipulating human language.
NLP subdomain focused on extracting meaning, intent and structure from text.
NLP subdomain focused on producing coherent and relevant natural-language text.
Family of autoregressive Transformer language models developed by OpenAI; GPT stands for Generative Pre-trained Transformer.
Language model with a very large number of parameters, typically trained on broad corpora at scale.
Language model with fewer parameters than LLMs, suitable for local or constrained deployment.
Model specifically fine-tuned to follow natural-language instructions and dialogue.
Pretrained model before any instruction or preference tuning; raw next-token predictor.
Conversational system that interacts with a user in natural language, often powered by an LLM.
Conversational AI system that helps a user accomplish tasks through natural-language interaction.
Special message at the start of a conversation that sets the model's persona, instructions and behavioral constraints.
Message from the user to the model in a conversation.
Message produced by the model in a conversation.
One exchange in a dialogue, consisting of a user message and an assistant response.
Sampling parameter scaling the logits before softmax; higher values produce more random outputs.
Sampling restricted to the k most probable next tokens; limits low-probability choices.
Sampling from the smallest set of tokens whose cumulative probability exceeds p, adapting the candidate set to the distribution.
Decoding strategy that always picks the most probable next token; deterministic but prone to repetition.
Decoding strategy keeping the k most probable partial sequences at each step to find high-probability outputs.
Decoding parameter penalizing tokens that have already appeared, to reduce repetitive output.
Integer initializing the random number generator to make experiments reproducible.
Maximum number of tokens a model is allowed to generate for a given response.
Throughput metric measuring the number of tokens generated or processed per second.
Latency between sending a request and receiving the first generated token; critical for interactive use.
Mode where the model returns tokens incrementally as they are generated, instead of waiting for the full response.
Decoding constrained to produce output that matches a given schema (JSON, grammar, regex, etc.).
Model output formatted as structured data such as JSON, validated against a schema.
Capability by which a model emits structured calls to external functions or tools in a predefined schema.
Generation of content that sounds plausible but is factually incorrect or unsupported by the input.
21 terms
Input provided to a language model to guide its behavior and generate a response.
Discipline of designing and optimizing prompts to obtain more accurate, reliable or controlled outputs from a model.
Prompting that asks the model to perform a task without providing any example.
Prompting that includes a single example of the desired behavior.
Prompting that includes a small number of examples to steer the model's behavior on a task.
Prompting technique that elicits intermediate reasoning steps before the final answer, improving performance on complex tasks.
Instruction in a prompt asking the model to show its reasoning step by step before concluding.
Breaking a complex task into smaller sub-tasks, each handled by a focused prompt or step.
Hidden or privileged prompt setting the model's role and global instructions, separate from the user prompt.
Prompt entered by the end user for a given request.
Parameterized prompt structure with placeholders for variables, enabling reuse across requests.
Placeholder in a prompt template filled with a value at runtime (user name, context, etc.).
Information included in the prompt (documents, retrieved passages, etc.) that the model should use to answer.
Marker in a prompt separating instructions from user-provided content, helping the model avoid prompt injection.
Technique chaining several prompts where each step's output feeds the next, to structure complex workflows.
Practice where the model is asked to critique its own answer before finalizing or returning it.
Workflow where a model proposes a draft and a human or another model iteratively revises it.
Example of an undesired behavior shown to the model in a prompt to discourage that behavior.
Role or character assigned to the model in a prompt to specialize its tone and style.
Instructions in a prompt that conflict with each other or with the model's prior training, potentially leading to inconsistent behavior.
Degradation of model performance when the context window is filled with irrelevant or excessive information.
26 terms
Architecture that retrieves relevant documents and includes them in the prompt to ground the model's answers in external knowledge.
Field concerned with finding relevant documents from a collection in response to a query.
Data structure enabling efficient search of documents or vectors matching a query.
Index mapping each term to the documents containing it; backbone of lexical search engines.
Search based on exact term matching (BM25, TF-IDF), without semantic understanding.
Classic ranking function for lexical information retrieval, weighting term frequency and document length.
Search based on meaning similarity rather than exact term matching, typically using embeddings.
Search by nearest-neighbor lookup of embeddings in a vector space.
Database optimized for storing embeddings and performing fast approximate nearest-neighbor queries.
Approximate Nearest Neighbor index trading exactness for speed at scale (HNSW, IVF, etc.).
Hierarchical Navigable Small World: graph-based ANN index offering fast queries with high recall.
Combination of lexical search (BM25) and vector search to benefit from both exact term matching and semantic similarity.
Contiguous segment of a document used as a retrieval unit in RAG pipelines.
Splitting of documents into chunks suitable for retrieval and model context.
Shared tokens between consecutive chunks to preserve context at boundaries.
Component of a RAG system that selects relevant documents or chunks for a query.
Second-pass reordering of retrieved candidates using a more accurate (and often more expensive) model.
Reranker that scores (query, document) pairs jointly, typically more accurate but slower than bi-encoders.
Transformation of the original query to improve retrieval (reformulation, expansion, decomposition).
Adding related terms or synonyms to the query to improve recall.
Restricting retrieval to documents matching certain metadata criteria (date, source, type, etc.).
RAG variant where the model links each claim to the source document and chunk supporting it.
Constraining model outputs to information actually present in the provided context, reducing hallucination.
User-owned document collection indexed for local RAG, without sending data to third parties.
Structured representation of entities and their relationships, used for reasoning and retrieval.
RAG variant leveraging a knowledge graph to retrieve structured, relation-aware context.
25 terms
System that uses a model to plan and execute actions via tools to accomplish a goal with limited human intervention.
Iterative loop where an agent observes, reasons, acts and re-observes until the goal is reached.
Component that coordinates the steps, agents and tools in an AI workflow.
Process by which an agent decomposes a goal into ordered steps before execution.
External capability (search, code execution, API call) that an agent can invoke to extend its actions.
Calling and running a tool from an agent, including handling inputs, outputs and errors.
Catalog of available tools with their schemas, descriptions and permissions.
Authorization policy defining which tools an agent may invoke and under which constraints.
Security principle granting an agent only the minimum permissions required to perform its task.
Workflow where a human validates or approves certain agent decisions before execution.
Checkpoint requiring explicit human approval before the agent performs a sensitive action.
Decision point in an agentic workflow where human or system validation is required before continuing.
Agent equipped with several tools that it can choose among depending on the task.
System in which several specialized agents collaborate, often under a supervisor, to accomplish complex tasks.
Agent that orchestrates other agents by assigning sub-tasks and aggregating their results.
Mechanism allowing an agent to retain context across turns or sessions, including conversation history and learned facts.
In-conversation context that an agent uses within a single session, typically bounded by the context window.
Memory that survives across sessions, often backed by an external store.
Snapshot of an agent's progress, including variables and intermediate results, enabling resumption.
Mechanism by which an agent detects, handles and recovers from tool or runtime failures.
Property of an operation that produces the same result whether run once or many times; critical for reliable agent tool calls.
Criterion determining when an agent should end its loop (goal reached, max iterations, error, etc.).
Isolated environment where an agent runs code or commands with restricted access to the host system.
Standard describing how a model or agent connects to and invokes external tools (e.g., MCP).
Open protocol standardizing how models and agents connect to data sources and tools.
25 terms
Field of AI that enables computers to interpret images and videos.
Task of assigning a category to an image from a predefined set of classes.
Task of locating and classifying objects in an image with bounding boxes.
Rectangle defined by coordinates that localizes an object in an image.
Task of assigning a class label to each pixel of an image.
Task of detecting and outlining each individual object instance in an image.
Task combining semantic and instance segmentation to label every pixel while distinguishing object instances.
Conversion of text in images into machine-readable characters.
Recognition of handwritten text from images or scans.
Model jointly processing images and text, capable of tasks such as captioning, VQA and visual reasoning.
Component of a multimodal model that encodes images into a representation usable by the language part.
Transformer architecture applied to images by splitting them into patches and treating each as a token.
Tile of an image used as input unit by a Vision Transformer.
Generation of images from a textual prompt, typically with a diffusion model.
Generation of a new image from a source image and a conditioning signal (prompt, sketch, etc.).
Filling in a masked region of an image with plausible content coherent with the surrounding area.
Extending an image beyond its original borders by generating new content consistent with the existing scene.
Generation of a high-resolution image from a low-resolution input.
Generation of video clips from text, images or other conditioning signals.
Consistency of objects, style and motion across frames of a generated video.
Task of following a specific object across a sequence of frames.
Task of detecting the position and orientation of bodies or objects from images or video.
Estimation of scene depth from a single image.
Implicit neural representation of a 3D scene that enables novel view synthesis by querying 5D radiance along rays.
Task of building a 3D model of a scene or object from images, video or sensor data.
25 terms
Generative model that learns to reverse a gradual noising process, producing high-quality samples.
Process that gradually adds Gaussian noise to data until it becomes pure noise; the reverse process is learned by the model.
Step where a model removes noise from a sample, iteratively denoising towards a clean image.
Diffusion performed in a compressed latent space rather than in pixel space, enabling efficient high-resolution generation.
Component controlling how noise levels evolve across denoising steps; affects quality and speed.
Algorithm that solves the reverse diffusion process (DDPM, DDIM, Euler, DPM++ etc.) to produce samples.
Number of denoising iterations used to generate a sample; a trade-off between speed and quality.
Classifier-Free Guidance scale: strength of conditioning applied during diffusion sampling.
Prompt describing what should be avoided in the generated output; used in diffusion models.
Additional signal (text, image, class) guiding the generation toward a specific output.
Neural structure that adds spatial conditioning (edges, depth, pose) to a pretrained diffusion model without retraining it.
Module connecting a pretrained diffusion model to an image encoder to provide visual conditioning.
Generation of music conditioned on an existing musical input (style transfer, accompaniment, etc.).
Synthesis of natural-sounding speech from text.
Transcription of spoken audio into text.
End-to-end conversion of speech from one voice or language to another, including prosody and identity.
Synthesis of a voice that mimics a specific speaker from a small audio sample.
Decomposition of a mixed audio signal into its constituent sources (speakers, instruments, etc.).
Segmentation of an audio recording by speaker, identifying 'who spoke when'.
Model jointly processing audio and text, enabling tasks such as audio captioning and audio QA.
Generation of video clips directly from textual prompts.
Generation of a video from a single image, animating the scene consistent with the initial frame.
AI-generated visual representation of a person, often driven by audio or text.
Alignment of mouth movements with synthesized speech for a realistic visual output.
Realistic synthetic media (image, audio, video) generated by AI to imitate a real person; raises ethical and legal concerns.
36 terms
Process of measuring a model's performance against metrics and benchmarks.
Standardized dataset and protocol used to compare models on a given task.
Quantitative measure used to evaluate a model (accuracy, F1, perplexity, etc.).
Fraction of correct predictions among all predictions; appropriate when classes are balanced.
Fraction of items predicted as positive that are actually positive.
Fraction of actual positive items that are correctly retrieved by the model.
Harmonic mean of precision and recall, balancing both concerns.
Table showing counts of true/false positives and negatives, used to analyze classification errors.
Case correctly predicted as positive by the model.
Case incorrectly predicted as positive; type I error.
Case incorrectly predicted as negative; type II error.
Case correctly predicted as negative by the model.
Curve plotting true positive rate against false positive rate at various classification thresholds.
Area Under the ROC Curve: aggregate measure of classification performance across thresholds; 1.0 is perfect, 0.5 is random.
Property that a model's predicted probabilities match observed frequencies of outcomes.
Metric measuring how well a language model predicts a sample; lower perplexity means better predictions.
Bilingual Evaluation Understudy: metric for machine translation based on n-gram overlap with reference translations.
Recall-Oriented Understudy for Gisting Evaluation: family of metrics for summarization based on n-gram recall.
Edit-distance-based metric for ASR systems, measuring the proportion of incorrectly transcribed words.
Strict metric: prediction is counted correct only if it exactly matches the reference answer.
Assessment of model outputs by human raters, often considered the gold standard for open-ended generation.
Evaluation in which raters do not know which model produced each output, reducing bias.
Evaluation where raters compare two outputs and pick the preferred one; foundation of preference data.
Evaluation where a strong LLM rates or compares model outputs as a scalable proxy for human judgment.
Proportion of generated responses containing hallucinated content, as judged by humans or automated checks.
Assessment of a RAG system's retrieval accuracy, answer faithfulness and citation quality.
Fraction of queries for which at least one relevant document appears in the top k retrieved.
Fraction of relevant documents among the top k retrieved.
Average across queries of the reciprocal rank of the first relevant document.
Normalized Discounted Cumulative Gain: ranking metric rewarding relevant items appearing early in results.
Ability of a model to maintain its performance under noisy, adversarial or out-of-distribution inputs.
Evaluation on data drawn from a different distribution than the training data, to assess generalization.
Test ensuring that new model versions do not degrade performance on previously solved tasks.
Online experiment comparing two model variants by randomly assigning them to users and measuring a target metric.
Range around an estimated metric that likely contains the true value with a given probability.
Probability that an observed difference between two systems is not due to chance.
31 terms
Discipline protecting AI systems from attacks such as prompt injection, data poisoning and model theft.
Property of a system behaving reliably and predictably, including failure modes, robustness and oversight.
Structured adversarial testing by experts attempting to elicit unsafe, biased or harmful behavior from a model.
Attack in which untrusted content smuggles instructions into a model's prompt, overriding the original instructions.
Prompt injection coming from the user's input directly (e.g., in a chat message).
Prompt injection smuggled via external content the model retrieves or processes (web pages, documents, emails).
Technique that bypasses a model's safety guardrails to elicit behavior it was trained to refuse.
Attack that extracts sensitive data (training data, user data, system prompts) from a model or its environment.
Attack that manipulates training data to introduce backdoors or biases into a model.
Attack that directly manipulates a model's parameters or training process to alter its behavior.
Hidden behavior triggered by a specific trigger, often inserted via data or model poisoning.
Crafted input designed to fool a model into misclassification or undesired output.
Input slightly perturbed to cause a model to err while remaining semantically identical to humans.
Attack that reconstructs a proprietary model by querying it with carefully chosen inputs.
Attack that recovers training data from a model's parameters or outputs.
Attack determining whether a specific example was part of a model's training set.
Tendency of models to reproduce training examples verbatim, raising privacy and copyright concerns.
Unintended disclosure of system instructions or hidden prompts to the end user.
Boundary separating components, data or users that can be trusted from those that cannot, used to scope validation.
Security strategy layering multiple independent controls (input filtering, output filtering, monitoring, HITL).
Validation and sanitization of inputs before they reach the model, blocking injections and unsafe content.
Inspection and validation of model outputs before they reach the user or downstream system.
Rule or component enforcing safety or behavioral constraints on model inputs and outputs.
Strict validation of the arguments passed to a tool by an agent, against a schema, to prevent misuse.
Explicit list of permitted values (tools, domains, sources), as opposed to a denylist; preferred for security.
Throttling of requests per user or per key to prevent abuse and protect availability.
Tamper-resistant record of significant events (requests, tool calls, overrides) for traceability and review.
Sum of entry points (APIs, tools, document loaders) through which a system can be attacked.
Set of components, vendors and processes that produce, train and distribute a model; integrity must be verified.
Practices ensuring model weights are not tampered with, exfiltrated or substituted between training and deployment.
Misuse of model compute or quotas for purposes other than intended (cryptomining, bulk scraping, etc.).
22 terms
Systematic skew in model outputs that disadvantages certain groups or reinforces social inequities.
Bias arising when training data is not representative of the population on which the model will be used.
Bias arising from under- or over-representation of certain groups in the training data.
Bias introduced by subjective or inconsistent choices of human annotators.
Tendency of humans to over-trust automated decisions, even when they are wrong.
Property of a model to treat individuals and groups equitably, according to a chosen formal criterion.
Outcome in which an AI system disadvantages individuals based on protected characteristics.
Ability to provide human-understandable explanations of a model's behavior.
Degree to which a human can understand the cause of a model's decision.
Explanation of a single prediction, identifying the features that drove it (e.g., SHAP, LIME).
Explanation describing a model's overall behavior across the dataset.
SHapley Additive exPlanations: method attributing a prediction to each feature based on Shapley values.
Local Interpretable Model-agnostic Explanations: method that locally approximates a model with an interpretable one.
Score quantifying the contribution of each feature to a model's predictions.
Minimal change to an input that flips the model's prediction; used to explain decisions and test robustness.
Property of a system to clearly document its capabilities, limitations, training data and intended uses.
Ability to reconstruct the chain of events, data and decisions that produced a given output.
Principle that a human remains responsible for decisions taken with the assistance of an AI system.
Mechanisms ensuring that humans can monitor, intervene and override an AI system's behavior.
Energy and resource footprint of AI systems, especially training and inference of large models.
Greenhouse gas emissions associated with the lifecycle of an AI system.
Situation where efficiency gains from AI lead to higher overall consumption, offsetting environmental benefits.
25 terms
Any information relating to an identified or identifiable natural person; protected by regulations such as the GDPR.
Irreversible removal of identifying information from data, so that re-identification is no longer possible.
Replacement of identifying data with artificial identifiers; reversible with additional information, hence not anonymization.
Principle limiting the collection and processing of personal data to what is strictly necessary for the purpose.
Specific, explicit and legitimate purpose for which personal data is processed; a foundation of data protection law.
Freely given, specific, informed and unambiguous agreement to the processing of personal data.
General Data Protection Regulation: European Union regulation governing the processing of personal data.
European Union regulation on artificial intelligence, adopting a risk-based approach to AI systems and models.
Regulatory principle imposing obligations proportional to the risks an AI system poses to fundamental rights and safety.
AI system classified as high-risk by the AI Act, subject to strict requirements (data governance, documentation, oversight).
Use of AI explicitly banned by regulation, such as social scoring or untargeted face scraping.
Requirement to inform users when they interact with an AI system, especially for generative content.
Analysis of the potential effects of an AI system on rights, safety and society before deployment.
Set of policies, processes and controls directing the development and use of AI within an organization.
Process of identifying, assessing, mitigating and monitoring risks associated with AI systems.
Structured information describing a model: intended use, training data, limitations, evaluations, etc.
Structured document describing a model's capabilities, intended use, limitations and evaluation results.
Structured document describing a dataset's motivation, composition, collection, preprocessing and recommended uses.
Legal terms under which a model can be used, modified and redistributed (commercial, research, open weights, etc.).
Legal terms under which a dataset can be used, especially for training AI systems.
Software whose source code is publicly available under a license allowing use, modification and redistribution.
Model whose trained weights are publicly released, allowing local inference and fine-tuning, under a chosen license.
Model whose weights and sometimes architecture are not publicly disclosed; usable only via an API or licensed product.
Set of legal questions on the protection of AI-generated works and the use of copyrighted data for training.
Mechanism allowing rights holders to request that their content not be used for training future models.
30 terms
AI running locally on the user's hardware (phone, laptop, edge device), without sending data to remote servers.
AI system that operates without network connectivity, typically on local models.
Deployment of an AI model on one's own infrastructure, retaining control over data and compute.
Running model inference on general-purpose processors, without requiring a GPU.
Running model inference on graphics processors, which offer large speedups for matrix operations.
Neural Processing Unit: specialized hardware accelerator optimized for neural network inference on edge devices.
Video RAM: dedicated memory on a GPU, limiting model size and batch size at inference and training time.
Main memory of the host system, used to host models too large to fit in VRAM.
Rate at which data can be read from or written to memory; a key bottleneck for large-model inference.
Technique reducing numerical precision of model weights and activations to lower memory and accelerate inference.
Quantization using 8-bit integer representations; common compromise between accuracy and efficiency.
Aggressive quantization using 4 bits per weight; widely used to run large models on consumer GPUs.
Quantization applied after training, without retraining, to reduce model size and speed up inference.
Training that simulates quantization effects during training to preserve accuracy after quantization.
Removal of redundant weights or neurons to reduce model size and compute with minimal accuracy loss.
Property of a model in which many weights are zero, enabling sparse computations and smaller storage.
Distribution of model weights or computation across heterogeneous devices (CPU RAM, disk, GPU) when the model does not fit on a single one.
Loading only a subset of model layers into GPU memory at a time, swapping as needed.
File format for storing quantized model weights, used by llama.cpp and compatible runtimes.
Tensor storage format designed for safe and fast loading of model weights, replacing older pickle-based formats.
Open Neural Network Exchange: interoperable format for representing trained models across frameworks and runtimes.
Open-source C++ runtime for efficient inference of LLMs on CPU and consumer GPUs, including quantized models.
Service exposing a model via an API on local infrastructure, for self-hosted LLM or ML deployments.
Software component that executes a model on hardware, handling memory, batching and optimizations (vLLM, TGI, llama.cpp, etc.).
Memory footprint of a model, typically measured in bytes or parameters.
Total number of trainable weights in a model; common indicator of capacity (millions, billions of parameters).
Memory consumed by the key-value cache during autoregressive generation, growing with sequence length.
Number of tokens (or requests) processed per unit of time by a serving system.
Time between sending a request and receiving the model's response.
Grouping multiple requests together to improve GPU utilization and throughput.
26 terms
Practices combining machine learning, DevOps and data engineering to deploy and maintain ML systems in production.
Adaptation of MLOps to LLM applications: prompt management, evaluation, monitoring, cost control, etc.
Sequence of automated steps that transform raw data into a trained and validated model.
Sequence of steps (preprocessing, model call, postprocessing) that turns an input into a final output at inference time.
HTTP or gRPC endpoint exposing a model for inference, usually with usage-based billing.
Network address exposing a service, such as a model inference API.
Process of making a model available in a target environment (cloud, on-premises, edge) for production use.
Infrastructure and software responsible for hosting a model and handling inference requests at scale.
Standardized unit packaging an application and its dependencies for portable, reproducible execution.
Automated management of containerized services: deployment, scaling, networking, healing (Kubernetes, etc.).
Mechanism allowing multiple models or tenants to share the same physical GPU, improving utilization and cost.
Automatic adjustment of compute resources based on load, to handle traffic variations while controlling cost.
Adding more instances of a service to handle load, as opposed to vertical scaling.
Increasing the resources (CPU, RAM, GPU) of a single instance, as opposed to horizontal scaling.
Ability to understand a system's internal state from its external outputs: metrics, logs and traces.
Tracking of a single request across multiple services to understand latency and failures end-to-end.
Recording of structured events emitted by a system for debugging, auditing and analysis.
Continuous tracking of a deployed model's performance, data drift and operational metrics.
Tracking of model versions, including weights, code and configuration, to ensure reproducibility.
Centralized catalog of trained model versions, with metadata, lineage and lifecycle states.
Ability to reproduce an experiment or training run bit-for-bit given the same code, data and environment.
Reverting a deployment to a previous, known-good version in case of regression.
Release strategy where a new version is gradually rolled out to a small fraction of traffic before full deployment.
Contract defining expected service levels (latency, availability, support) between provider and customer.
Limit on the number of tokens a user, team or feature can consume over a given period, to control cost.
Monetary cost of processing a single inference request, typically combining token and infrastructure costs.
16 terms
System that suggests items (products, content, contacts) to users based on their profile and behavior.
Recommendation approach based on patterns of similar users or items, without explicit content analysis.
Search whose results are reordered based on the user's profile, history or preferences.
Identification of observations that deviate significantly from the expected pattern (fraud, failures, outliers).
Anticipation of equipment failures using models trained on sensor data and historical incidents.
Prediction of future values of a series from its past values, often with statistical or deep models.
NLP task of determining the polarity (positive, negative, neutral) or finer emotions of a text.
NLP task of identifying and classifying named entities in text (people, organizations, locations, etc.).
Generation of a concise summary preserving the essential information of a source document.
Automatic translation of text from a source language to a target language.
Generation of source code from a specification, typically a natural-language prompt.
Suggestion of the next tokens of code as a developer types, based on context and prior code.
Use of an AI model to analyze a diff and suggest improvements, bugs or style issues.
Identification of fraudulent transactions or behaviors, often via anomaly detection and classification models.
Use of AI for perception, planning and control in robots operating in the real world.
Virtual replica of a physical system used for simulation, monitoring and optimization.
19 terms
Behavior that looks like understanding to a human observer, without implying genuine comprehension by the model.
Hypothetical capacity of an AI system for subjective experience; remains speculative and unestablished.
Tendency to attribute human traits, emotions or intentions to AI systems, which can mislead users.
Degree to which a model's statements are factually correct, often measured against trusted sources.
Property that a model's expressed confidence matches its actual accuracy.
Uncertainty due to limited knowledge or data, reducible by gathering more information.
Irreducible uncertainty inherent to the data, even with perfect knowledge.
Statistical association between two variables; does not imply causation.
Relation where one variable directly influences another, going beyond statistical correlation.
Reasoning about cause-and-effect relationships, often via causal models or counterfactuals.
Practice of optimizing specifically for benchmark scores without improving real-world capabilities.
Principle stating that 'when a measure becomes a target, it ceases to be a good measure'.
Property of a system producing the same output for the same input, given a fixed random seed.
Property of a system that can produce different outputs for the same input, due to sampling or parallel non-associative operations.
Joint tradeoff between output quality, monetary cost and response time that must be balanced in production.
Date after which a model has no reliable knowledge of world events, due to the end of its training data.
Mechanism for refreshing a model's knowledge (RAG, fine-tuning, retraining) after its cutoff date.
Reference invented by a model that does not correspond to a real source, a frequent hallucination pattern.
Mandatory human check of AI outputs before they are acted upon, especially for high-stakes use cases.