Tensor Cores and Matrix Units Explained for Chip Designers (2026)

A tensor core is a specialised execution unit built into a chip that performs matrix multiply-accumulate — written as D = A x B + C — on a whole small tile of operands in a single instruction, at many times the throughput of a general-purpose core for that one operation. A matrix unit is the hardware designer term for the same idea, and the two words describe slightly different emphases. This guide covers how the hardware works, how the architectures differ across vendors, and what the constraints are that decide whether a real kernel gets any of the speed on paper.

Table of Contents

How Tensor Cores and Matrix Units Handle AI Workloads

How Tensor Cores and Matrix Units Handle AI Workloads

Almost every arithmetic-heavy layer in a modern neural network reduces to the same primitive: multiply two matrices together. Attention projections, feed-forward layers, the batched matrix operations in a decoder, the convolution turned into an im2col GEMM — all of it ends up as dense matrix multiplication, usually called GEMM after the BLAS routine that does it.

That is the entire reason matrix hardware exists. A general-purpose core executes one instruction per thread per clock and spends most of its energy on instruction decode, register movement and address arithmetic. A matrix unit throws all of that away: it fetches two fragments into registers, runs a fixed grid of multipliers over them, and writes back a tile of results. Nothing else happens in the datapath.

The payoff shows up in two very different places depending on the workload. In training, and in the prefill phase of inference where a whole prompt is processed at once, you are doing arithmetic intensity and compute is the bottleneck. The matrix unit is the bottleneck resource, so widening it moves the wall clock time.

In the decode phase of inference, where the model emits one token at a time, you are reading a model’s worth of weights out of memory for a handful of matrix operations. There, memory bandwidth sets the pace and the matrix unit is mostly idle. Batching requests is what pushes decode back toward compute-bound, which is why the same chip behaves so differently under a busy serving load than a one-off benchmark.

Everything else these units touch — convolutions, sparse matrix multiply, FFT and Hadamard transforms, prefix scans, reductions, dynamic programming — is either GEMM in disguise or a reduction that a fixed-size array happens to do well.

What Is a Tensor Core?

A tensor core is a fixed-function unit that computes one small matrix multiply-accumulate per instruction, cooperatively across a group of threads. The three operands have fixed roles: A and B are the input fragments, C is the accumulator read in and added, D is the result written out.

Tensor cores and matrix units explained in one instruction

On NVIDIA hardware the instruction names the whole geometry, for example mma.sync.aligned.m16n16k16.row.col.f32.f16.f16.f32. Read that as: synchronous, warp-collective, 16 rows by 16 columns of output accumulated over a depth of 16, with FP16 inputs and an FP32 accumulator. Ampere added 8×8 forms, Hopper introduced larger warp-group forms such as wgmma that are issued by a group of warps rather than one.

Nothing about that instruction is negotiable at run time. The tile shape is baked into the silicon, which is why matrix dimensions need to be multiples of 8 or 16 for the fast path, and why a matrix that does not fill the array wastes hardware.

The vocabulary trips people up more than the hardware does, so it is worth one short detour. A tensor is just a labelled array of numbers. A matrix is a rank-2 tensor, a vector a rank-1 tensor. When a neural network layer computes Y = XW, it is contracting a tensor dimension — multiplying and summing across the shared axis. “What is a tensor matrix” usually means exactly this: the object a tensor core consumes is a plain matrix or a stack of them, and a “tensor matrix” as a separate type of thing does not exist.

AspectTensor core (NVIDIA naming)CUDA / shader core
Unit typeFixed-function matrix datapathGeneral-purpose vector pipeline
OperationMatrix multiply-accumulate on a fixed tileAny scalar or vector instruction
Data typesFP16, BF16, TF32, FP8, FP4, INT8, INT4FP32, FP64, INT32, and the same narrow formats
ParallelismOne warp-collective or warp-group instructionOne thread per instruction
Control flowFixed sequence, no branches inside a tileFull branch, call and predicate support
Programmer viewThrough WMMA, mma.sync, or high-level librariesDirect, instruction by instruction
Best fitDense and structured-sparse GEMMEverything else, plus small or oddly shaped math

One recurring misconception deserves a line of its own. Tensor cores do accelerate FP32 work — through the TF32 path, which keeps a 10-bit mantissa inside a 32-bit format, and through any of the narrow formats cuBLAS and CUTLASS will hand you. A plain IEEE FP32 GEMM does not run on the tensor path at all, so the honest answer is “yes, with a different number format, and you have to ask for it”.

What Is a Matrix Unit?

A matrix unit is a block of silicon whose only job is multiplying and adding matrix elements. Where “tensor core” names the product feature, “matrix unit” names the design object, and it is how an architect describes the block on a block diagram.

The defining implementation choice is the systolic array. Instead of pulling every operand out of memory for every single multiply, you stream A rows in from one edge and B columns in from the top, and each partial product hops one grid cell per clock until it lands in the accumulator register file at the far corner. The rhythm of the data movement is what keeps the multipliers fed.

That dataflow has a real design payoff. Each operand is fetched once and reused across the whole array, so arithmetic intensity inside the block is high by construction and the array needs far less register file bandwidth than a traditional design. The cost is rigidity: fixed dataflow, fixed tile shape, and a hard requirement that dimensions align.

Register blocking is the companion technique. Each thread or warp holds a small patch of C in registers and loops over K, reloading a fresh A fragment and B fragment per step. Get the blocking right and the tensor path runs near its peak; get it wrong and you stall on operand fetch, which is the single most common reason a kernel issues tensor instructions and still runs at a fraction of the advertised rate.

For a design team, a matrix unit is also an IP decision. A hard macro gives you a proven array, a datapath and a verified model, but a fixed geometry and a vendor-tied interface. A soft IP version lets you set tile shape, precision mix and register width yourself, at the cost of the verification work. Most teams buying a matrix unit today are picking between those two, and the tile geometry they are stuck with for a product’s lifetime is the most expensive thing on the datasheet.

Tensor Cores vs. Matrix Units: What Changes?

Mostly the emphasis and the control model, not the physics. Both are matrix multiply-accumulate engines, and no implementation on the market is identical to any other. The differences that matter to a designer fall into a handful of buckets.

Terminology. “Tensor core” is NVIDIA’s brand for a feature and is now used loosely. AMD says Matrix Core, Google says MXU, Apple and Intel say AMX, FPGA vendors say DSP block. Underneath, all of them multiply matrices in a fixed array.

Granularity of control. NVIDIA’s units are driven by per-instruction fragments of a few dozen elements. Google’s MXU is a large array driven a whole tile at a time by compiled code, and Apple’s and Intel’s AMX are wide tiles loaded explicitly into a dedicated block of registers. The fine-grained model gives more scheduling freedom; the coarse one gives more predictable throughput per cycle and far fewer instructions to issue.

Register file depth. This is the quiet one. How many C elements a unit can hold while it accumulates determines how much of a K loop it can absorb before spilling. Deep accumulator files let a large tile run at full rate; shallow ones force smaller tiles and more memory traffic per FLOP.

Precision menu. A unit built for FP16 only will not accept INT8, and a unit built for INT8 only will not do FP8. Some designs add dedicated sparsity paths, some a second block for fine-grained transforms.

Placement. A GPU tensor core sits beside general-purpose cores and shares the memory hierarchy. An MXU sits next to a vector memory designed to feed it. An AMX block sits beside a CPU. Same operation, completely different surrounding system.

How Matrix Multiplication Becomes Hardware Work

The mapping is easier to see with numbers. Take a small C = A x B where A is 2×2 and B is 2×2: A row 1 is 1, 2; A row 2 is 3, 4. B column 1 is 5, 6; B column 2 is 7, 8.

C[0][0] is 1 times 5 plus 2 times 6, which is 5 plus 12, which is 17. C[0][1] is 1 times 7 plus 2 times 8, which is 7 plus 16, which is 23. C[1][0] is 3 times 5 plus 4 times 6, which is 15 plus 24, which is 39. C[1][1] is 3 times 7 plus 4 times 8, which is 21 plus 32, which is 53.

Four outputs, eight multiplies, four adds, and a loop over K that a general-purpose core would run with each element loaded, multiplied, added and stored individually. On a matrix unit, the A fragment and the B fragment are loaded once, eight partial products are computed in parallel, and the four sums accumulate into the C registers — the D = A x B + C form, in one pass.

Scale that up and the real problem appears. A 4096×4096 by 4096×4096 GEMM is roughly 137 billion multiply-accumulates, and the naive layout reads every input element from memory about 4096 times. A tiled implementation walks K in steps, holds an output tile in registers, and drops total memory traffic by orders of magnitude.

So the hardware sets the rhythm, and the tiling sets the traffic. Neither alone gets you the throughput.

How Do Tensor Cores and Matrix Units Improve AI Performance?

There are five factors, and they interact. Understanding which one is binding tells you why a kernel is or is not fast.

Arithmetic throughput per instruction. One tensor instruction replaces hundreds of scalar multiply-adds, which is where the headline numbers come from. It is also the least useful number on its own.

Memory movement. The roofline model puts a hard ceiling on any kernel: its achievable rate is the lower of the machine’s peak compute and peak memory bandwidth times the kernel’s arithmetic intensity in FLOPs per byte. A GEMM with well-chosen tiles has high arithmetic intensity and runs under the compute roofline. A badly tiled one drops under the bandwidth roofline and no amount of extra tensor throughput helps.

Precision. Halving the bytes per operand cuts memory traffic and often doubles the achievable multiply rate. This is why FP8 and FP4 moved the practical ceiling more than the raw FLOP figures suggest.

Sparsity. Structured 2:4 sparsity, where every group of four consecutive values has two zeros, can double effective throughput because the hardware skips the zero pairs. The requirement is rigid: a dense matrix gets nothing from it, and a lightly pruned one gets nothing either.

Utilisation. Occupancy, shared memory bandwidth, bank conflicts, async pipelining, warp specialisation, alignment and tile shape all decide how much of the hardware you actually keep busy. A GEMM that issues only mma.sync instructions and spends 70% of its time waiting on shared memory loads has a tensor path and no speed.

One more practical note: the “tensor core count” printed on a GPU spec sheet is a per-part number of fixed blocks, not a performance rating. Two chips with the same count differ widely once clock, memory bandwidth, cache size and tile geometry enter the picture, and the number stops predicting anything once clock limits or thermals take over.

Where Are They Used in GPUs, NPUs, and Other Chips?

Every segment that runs dense AI now ships a matrix unit. The name changes, the array geometry changes, and the surrounding system changes more than the core operation does.

Vendor and partUnit nameArray or tile shapeTypical precisionsControl modelReached through
NVIDIA, Volta onwardTensor CoreSmall per-warp tiles, for example m16n16k16, with larger warp-group tiles on HopperFP16, TF32, INT8, then BF16, FP8, FP4Warp-collective or warp-group instructionCUDA via cuBLAS, cuBLASLt, CUTLASS, WMMA, mma.sync, wgmma
AMD, RDNA and CDNAMatrix CoreWMMA-style tiles per workgroupFP16, BF16, INT8, FP8 on newer partsWorkgroup-level wavefront instructionROCm via rocBLAS, HIP
Google, TPUMXULarge 128×128 systolic arrayBF16 primary, INT8 and other formats per generationWhole-tile operation, no shape dualityXLA, Pallas, TorchTPU
Apple, M-series and A-seriesAMX512-byte tile across a block of registersINT8 and BF16Explicit tile load, then a single wide operationCore ML, oneDNN, Metal
Intel, Xeon and CoreAMX512-byte tile, 8 rows x 64 columns of BINT8 and BF16Explicit tile load and executeoneDNN, Intel Extension for PyTorch
FPGA vendorsDSP block or soft matrix IPConfigurable, often 4×4 or smaller with resource sharingFixed point, usually 18 or 27 bit multipliersFully programmable, winsor-style or vendor DSPVendor toolchains and HLS

Designers pick different shapes because the markets differ. A gaming GPU needs one matrix unit sitting beside thousands of general-purpose cores, tuned for small matrices arriving constantly. A TPU is built for one workload, so it can afford a giant array and a memory system that exists to feed it. An AMX block in a phone or a server CPU only takes a few percent of the die, so it trades peak rate for the ability to run a whole layer without leaving the core.

The performance questions people ask about TPUs versus GPUs are really questions about systems. An MXU is a larger array than a tensor core, but a TPU is a different memory hierarchy, a different compiler path and a different scaling story. Neither is better in the abstract.

Precision, Data Types, and Accuracy Tradeoffs

Every matrix unit supports a small menu of number formats, and each generation adds narrower ones. The critical distinction is between input precision and accumulator precision. You can feed FP16 values and accumulate into FP32 registers, which is the standard mixed-precision recipe: the inputs are small and fast, and the running sum stays wide enough not to lose the plot over a long K loop.

FormatArrived withTypical accumulateWhere it is usedMain risk
FP32BaselineFP32, general core onlyReference maths, non-GEMM codeNo tensor path at all
TF32AmpereFP32Training convergence experiments10-bit mantissa loses precision
FP16VoltaFP16 or FP32Training and inference, the workhorseNarrow range overflows without a scale factor
BF16AmpereFP32Training, where stability matters more than mantissa bitsCoarser than FP16
FP8Hopper, AdaFP16 or FP32Large-model inference and trainingPer-block scaling required
INT8VoltaINT32Quantised inferencePer-tensor calibration, accuracy loss
INT4Later partsINT32Heavily quantised weightsNoticeable quality drop
FP4 and FP6BlackwellFP16 or FP32Serving large models, memory-bound decodeAggressive; needs fine-grained scaling

Accumulator width is where deployments quietly go wrong. Sum a long K dimension in FP16 and you lose the small terms, no matter how clean the inputs are. That is why every serious training stack promotes the accumulator to FP32, and why the Transformer Engine exists as a library that manages scale factors, promotes and demotes formats per layer, and tracks numerical error statistics as it runs.

Switching to FP16 does not, by itself, activate a tensor core. The dimensions have to fit a supported tile, the instruction the compiler picks has to be a tensor instruction, and the library has to choose the tensor path. Software decides, not the data type alone.

How to Choose the Right Hardware for an AI Workload

Start from the workload, not the spec sheet. Five questions separate a good fit from an expensive mismatch.

What shape are your matrices? If they tile cleanly into the array’s geometry, matrix hardware wins big. If they are thin, ragged, or dominated by small operations, a general-purpose core may do better no matter what the datasheet says.

Compute-bound or memory-bound? Training and prefill are compute-bound, so peak matrix throughput and supported precision decide the answer. Token-by-token decode is memory-bound, so memory bandwidth and capacity come first, and the matrix unit is not the constraint you should be shopping on.

What latency and batch size? Single-request interactive decode rewards bandwidth and small kernels. Throughput serving rewards large batches, which is exactly the regime where batched GEMM and matrix units reach their best utilisation.

What is the power and thermal budget? Matrix units deliver far more FLOPs per watt than general-purpose cores. That advantage is largest in mobile and edge parts, where it is often the reason the design is viable at all.

Will the software path exist? A chip with a world-class array and no compiler support for your model shape is a slower chip. Check that your framework, your quantisation scheme and your target batch sizes have tuned kernels before you commit. This is the shortest answer to “is a TPU better than a GPU”: neither, unless one of them has a library that matches your model.

For an IP buyer, add two more. Confirm whether the tile geometry is fixed or configurable, because you cannot change it after tape-out. And confirm which precisions have a dedicated path versus a conversion path, because the second option is much slower than the datasheet implies.

Tensor Cores and Matrix Units Explained: Key Takeaways

Tensor cores and matrix units explained, in one line: they are fixed-function matrix multiply-accumulate engines that compute a tile of D = A x B + C per instruction, and the name tells you the vendor, not the design.

The distinctions that matter in practice are four. Terminology varies but the physics does not. Granularity and accumulator depth differ, and they set how much of the array you can keep busy. Precision choice sets your memory traffic and your numerical risk, and input precision is not the same as accumulator precision. And throughput only materialises when the tile shape, alignment, layout and data supply all line up.

If you take one thing away, learn the D = A x B + C operand roles and one real instruction name for your target architecture. Everything else — the roofline, the fragments, the compiler flags — becomes much easier to reason about once that sits clearly in your head.

Frequently Asked Questions

Can you explain what tensor cores are and how they work?

A tensor core is a specialised execution unit that runs a matrix multiply-accumulate, D = A x B + C, on a small tile of operands in one instruction. Input fragments A and B are loaded into registers, a grid of multipliers computes every partial product in parallel, and the sums accumulate into C registers to produce D. The instruction is collective: a warp or warp group issues it together, and the tile geometry is fixed in hardware.

What is a tensor matrix?

There is no separate object called a tensor matrix. A tensor is just a labelled array of numbers, so a matrix is a rank-2 tensor and a stack of matrices is a rank-3 tensor. What a tensor core consumes is an ordinary matrix multiply, where a shared dimension is contracted: multiply along the shared axis and sum. Most people asking this question mean the GEMM operation that dominates neural network layers.

What are the key differences between CUDA cores and tensor cores?

A CUDA core is a general-purpose vector pipeline: one thread, one instruction, any operation, full branch support. A tensor core is fixed-function: one instruction performs a complete matrix multiply-accumulate across a warp, in a few number formats, with a tile shape baked into the silicon. Tensor cores win enormously on dense GEMM and are close to useless for control flow, irregular shapes or small operations.

Do tensor cores accelerate FP32?

Partly. A plain IEEE FP32 GEMM does not run on the tensor path, so it stays on general-purpose cores. But TF32, a 32-bit format with a 10-bit mantissa, runs on tensor cores, as do FP16, BF16, FP8 and INT8 with wider accumulators. The practical effect is that FP32-precision code can use tensor hardware, but only after you accept a reduced mantissa or switch formats, and only if the library you call offers that path.

How is matrix multiplication done on a GPU?

A GPU decomposes a GEMM into tiles. The output matrix is split into blocks that fit in registers, and K is walked in steps. For each step the hardware loads one A fragment and one B fragment from shared memory, issues a single collective matrix instruction, and accumulates the result into C registers held in registers across many steps. The tile size, the load pattern and the memory layout decide whether the tensor path stays saturated or stalls waiting on operands.

Does using FP16 automatically activate tensor cores?

No. Precision alone does nothing. Three things have to line up: the matrix dimensions must tile into a geometry the hardware supports, the compiler or library must choose a tensor instruction rather than a sequence of scalar ones, and the data must be laid out so fragments load without alignment penalties. A kernel that uses FP16 but tiles badly, or that the library has no tuned path for, can run slower than a plain FP32 kernel on general-purpose cores.

Leave a Comment