TFLOPS is FLOPS measured in units of $10^{12}$ (teraflops) floating-point operations per second, the scale at which modern GPUs and small clusters are quoted. A single high-end datacenter GPU exceeds tens of TFLOPS in FP64 and over a hundred TFLOPS in lower-precision formats used for machine learning.
Precision matters enormously: a GPU's FP16 tensor throughput can be 10x higher than its FP64 throughput, and quoting “TFLOPS” without precision is meaningless.
GPU TFLOPS scales through massive parallelism, not single-thread speed.
GPU TFLOPS: thousands of narrow lanes × lower precision CPU GFLOPS: handful of wide SIMD units × high precision