Tensor cores
Specialized matrix-multiply units on NVIDIA RTX GPUs, optimized for the half-precision math used in deep learning. Power DLSS, frame generation, and on-device AI inference.
Tensor cores execute fused multiply-accumulate operations on small matrices (typically 4×4 FP16/BF16/INT8) far faster than CUDA cores can. The architectural payoff: hundreds of TOPS on consumer cards.
What uses them
- DLSS — neural upscaling and frame generation.
- NVIDIA Broadcast — background blur, noise removal.
- Local LLMs — Llama, Mistral inference at usable speed.
- Stable Diffusion — image generation 5–10× faster than CUDA-only.
Generations
- Turing — 1st gen.
- Ampere — 3rd gen, added sparsity.
- Ada — 4th gen, added FP8.
- Blackwell — 5th gen, added FP4 + much higher throughput.
How to use Tensor cores in a real comparison
Specialized matrix-multiply units on NVIDIA RTX GPUs, optimized for the half-precision math used in deep learning. Power DLSS, frame generation, and on-device AI inference. In practice, this is most useful as one part of a decision rather than a standalone quality badge. Compare it alongside the workload, room, ecosystem, budget, and ownership constraints that apply to you; the strongest published figure is not automatically the best outcome for every buyer.
What to check next
Open a product page and check the stated value, configuration, price, and related trade-offs before treating this term as decisive. It is especially relevant in .
Avoid the one-number trap
Manufacturers can describe the same capability under different conditions. Compare like-for-like variants, check the unit and test condition where available, and use a head-to-head page to see whether a measurable difference is material for your own use.
vsMars presents structured catalog values to make trade-offs legible. The linked product and spec pages are the right place to inspect the underlying values before buying.