QBit
QBit functions reference.
dequantizeInt8ToBFloat16
Reconstructs a BFloat16 value from an Int8 code produced by
quantizeBFloat16ToInt8, by mapping the code back to the
reconstruction level of its Gaussian Lloyd-Max cell.
The argument may be a single Int8, an Array(Int8), or a QBit(Int8, ...); the function is
applied to every element and returns BFloat16, Array(BFloat16), or a QBit(BFloat16, ...) of
the same dimension and stride respectively.
Syntax
dequantizeInt8ToBFloat16(x)Arguments
x— Code(s) produced by quantizeBFloat16ToInt8.Int8orArray(Int8)orQBit(Int8)
Returned value
The reconstruction level of the cell. BFloat16 or Array(BFloat16) or QBit(BFloat16)
Examples
Round trip
SELECT round(toFloat32(dequantizeInt8ToBFloat16(quantizeBFloat16ToInt8(0.5::BFloat16))), 4)0.4961Introduced in version 26.7.0.
quantizeBFloat16ToInt8
Quantizes a BFloat16 value to Int8 using a 256-level Gaussian Lloyd-Max quantizer
(the MSE-optimal scalar quantizer for a standard-normal source).
Intended for values that are approximately distributed as N(0, 1). A random orthogonal
(randomized Hadamard) rotation preserves the norm, so for a d-dimensional unit-norm embedding
each rotated coordinate has variance 1/d; scale the rotated vector by sqrt(d) to reach unit
variance before quantizing.
The argument may be a single BFloat16, an Array(BFloat16), or a QBit(BFloat16, ...); the
function is applied to every element and returns Int8, Array(Int8), or a QBit(Int8, ...) of
the same dimension and stride respectively. Passing a whole vector (Array or QBit) is preferred
over arrayMap(x -> quantizeBFloat16ToInt8(x), ...).
The result is the index 0..255 of the Lloyd-Max cell, stored as Int8 as index - 128, so
the sign bit equals the sign of the value (+0/-0 map to the positive/negative central cells)
and the top b bits form a valid embedded 2^b-level quantizer: bit-truncation of the code
yields coarser Int4/Int2/binary codes.
Use dequantizeInt8ToBFloat16 to reconstruct the value.
Syntax
quantizeBFloat16ToInt8(x)Arguments
x— Value(s) to quantize (expected to be ~N(0,1)).BFloat16orArray(BFloat16)orQBit(BFloat16)
Returned value
The Lloyd-Max cell index minus 128. Int8 or Array(Int8) or QBit(Int8)
Examples
Quantize a scalar
SELECT quantizeBFloat16ToInt8(2.0::BFloat16)116Quantize an array
SELECT quantizeBFloat16ToInt8([0.1, -0.5, 2.0]::Array(BFloat16))[10,-49,116]Introduced in version 26.7.0.