AI Glossary

The AI terms that
actually matter.

Clear, technical definitions of 115 key concepts — written by AI engineers, not marketers.

AI Architecture

Mamba Model

The Mamba model is a groundbreaking AI architecture that achieves state-of-the-art performance across various modalities by utilizing selective state spaces for efficient, linear-time sequence modeling

Read definition
AI Techniques

Matryoshka Embeddings

Matryoshka embeddings are vector embeddings trained so that a short prefix of the full vector (say, the first 128 of 1536 dimensions) is itself a usable, accurate embedding, letting one model serve many storage and latency budgets instead of needing a separate model per size.

Read definition
AI Models

MiniMax M2.7

MiniMax M2.7 is an open-source, self-evolving Mixture-of-Experts AI model that autonomously participates in 30–50% of its own training workflow, featuring 230 billion total parameters, a 200K-token context window, and top-tier performance on agentic and software engineering benchmarks.

Read definition
AI Architecture

Mixture of Depths (MoD)

Mixture of Depths is a transformer routing technique where a learned per-layer router selects only a fixed fraction of tokens to pass through that layer's full attention and MLP block; the rest skip the block entirely via a residual connection. Because the fraction (capacity) is fixed ahead of time, total compute stays static and predictable, unlike Mixture-of-Experts routing.

Read definition
AI Architecture

Mixture of Experts (MoE)

Mixture of Experts is a neural network architecture where a learned routing mechanism activates only a small subset of specialised sub-networks (experts) for each input token — delivering the capacity of a much larger model at a fraction of the per-token compute cost.

Read definition
AI Techniques

Model Context Protocol (MCP)

Model Context Protocol (MCP) is an open standard introduced by Anthropic that defines a universal interface for connecting AI models to external data sources, tools, and services — eliminating the need for bespoke integrations for every AI-to-system connection.

Read definition
AI Techniques

Model Merging

Model merging combines the weights of two or more independently trained or fine-tuned neural networks into a single model, without backpropagation or extra training, producing a model that inherits multiple source models' capabilities at zero added inference cost.

Read definition
AI Models

Moshi

Moshi is a full-duplex, real-time spoken dialogue AI developed by Kyutai that can simultaneously listen and speak — enabling natural, interruption-aware conversations with theoretical latency as low as 160ms.

Read definition
Agentic AI

Multi Agent System

A Multi-Agent System (MAS) is an AI architecture where multiple autonomous, specialized agents interact, collaborate, or compete to solve complex workflows that are too difficult or broad for a single, monolithic AI model.

Read definition
Multimodal Models

Natively Multimodal Model

A natively multimodal model is trained from scratch on multiple modalities at once, text, images, audio, video, rather than bolting a frozen vision or audio encoder onto a finished language model. The modalities share one token stream and one set of weights, which is what 'native' refers to.

Read definition
AI Architecture

NVIDIA Cosmos 3 Edge

NVIDIA Cosmos 3 Edge is a 4-billion-parameter open world foundation model, the smallest tier of the Cosmos 3 family, that reasons over vision and generates robot actions directly on Jetson-class edge hardware without a cloud round-trip.

Read definition
ASR

NVIDIA Parakeet TDT

NVIDIA's Parakeet TDT 1.1B is the fastest open-source ASR model available, achieving an RTFx near 2,000× real-time — processing audio 6.5× faster than Canary-Qwen at the cost of some accuracy.

Read definition
← Previous Page 6 of 10 Next →

Stay ahead of the curve

Weekly newsletter on agentic AI, LLMs, and what we're building at Superteams — straight to your inbox.

Ready to ship AI in production?

We deploy fractional AI teams that deliver production-grade systems in 30–90 days. No fluff, no obligation.

Book a strategy call