Dekum: Tapered Precision Decimal Floating-Point Format
Abstract
This document introduces the ‘dekum’ format, a tapered-precision decimal floating-point representation. The currently established standard for decimal floating-point arithmetic is IEEE 754-2008, which specifies two alternative methods for encoding decimal numbers. By design, a substantial fraction of these representations are redundant, and the formats are fixed to specific bit widths (32, 64, and 128 bits). In contrast, a number format should ideally be definable for any bit width. Dekums are inspired by takums and constitute a tapered precision format with a bounded dynamic range. By design, they admit no redundant representations.
Full text
Dekum: Tapered Precision Decimal Floating-Point Format Laslo Hunhold 9th November 2025 1 introduction It is a well-known fact that the expression 0 . 1+0 . 2 cannot be represented exactly in binary floating-point arithmetic, yielding, for instance, a result of 0 . 2998 when using float16 . This circumstance lends binary floating-point arithmetic a certain air of unreliability, but it also has far-reaching consequences. Decimal notation is the natural human way of expressing numbers, and errors arising from inexact binary representations can accumulate over millions of computations, for example in financial applications. For this reason, decimal floatingpoint arithmetic has become the norm in such domains, and it is often expected to offer advantages over binary floating-point formats. This document introduces the ‘dekum’ format, a tapered-precision decimal floating-point representation. The currently established standard for decimal floating-point arithmetic is IEEE 754-2008, which specifies two alternative methods for encoding decimal numbers. By design, a substantial fraction of these representations are redundant, and the formats are fixed to specific bit widths (32,64, and 128 bits). In contrast, a number format should ideally be definable for any bit width. Dekums are inspired by takums and constitute a taperedprecision format with a bounded dynamic range. By design, they admit no redundant representations. A particular challenge of representing decimal floating-point numbers in binary form arises from the fundamental incompatibility between powers of ten and powers of two. A frequently used intuition is that 210 =1024 is close to 103=1000 , implying that ten binary digits can approximately encode three decimal digits. While this approximation is serviceable, it nevertheless wastes 24 binary states. When concatenating multiple ten-bit blocks to represent multiples of three decimal digits, combinatorial effects result in a significant number of wasted representations. Consider a sequence of m ten-bit blocks, each encoding three decimal digits. The total number of binary states is 210m , while 1
2 definition 2 sign exponent significand SDRCI 1 1 3 c a Figure 1: The dekum n -bit binary format for n⩾10 , with sign bit S ,direction bit D ,regime bits R ,characteristic bits C and significand bits I . The quantity c is the characteristic bit count and a:= n−c−5 is the accuracy. the total number of distinct decimal combinations is 103m . The number of wasted binary representations is therefore 210m −103m , and the ratio of wasted representations is given by 210m −103m 210m =1−(︃1000 1024)︃m . (1) For m∈{2 , 5 , 11} , corresponding to the sizes used in IEEE 754 decimal32 , decimal64 , and decimal128 , the ratios are approximately 5 %, 11 %, and 23 %, respectively. In addition, the explicit storage of trailing-zero variants of the same number (e.g. 1 , 1 . 0 , 1 . 00 ) introduces an additional waste of approximately 1 % to 11 %, compounded by redundant encodings for infinity and non-real values. Moreover, a considerable portion of the representable range remains practically unused due to excessive dynamic range. In summary, there remains substantial potential for improvement in the design of decimal floating-point formats, particularly with respect to representational efficiency. Such improvements may even help restore their practical relevance, given the well-documented shortcomings of binary floating-point arithmetic. 2 definition The dekum binary format is given in Figure 1without loss of generality for n⩾10. The regime value ris determined as r:= {︄uint(R)D=0 uint(R)D=1,(2) and the characteristic bit count follows as c:= max(0,r−2). (3) We define a bias b:= {︄(−1,−2,−3,−5,−9,−17,−33,−65)rD=0 (0,1,2,3,5,9,17,33)rD=1(4)
2 definition 3 and obtain the exponent value e:= b+uint(C)∈{−65,...,64}. (5) This yields a general-purpose dynamic range of approximately 10±65 , which is pretty ideal. We obtain the significand value d∈[1 , 10) by first determining the full set of possible decimal values with the given accuracy a using the following scheme: One first computes K⩾−1 and the residual R∈N0such that 2a=9+(︄K ∑︂ k=0 9·10k·9)︄+R=9+81 (︄K ∑︂ k=0 10k)︄+R. (6) This observation indicates that K+2 decimal digits can be fully and uniquely represented, with a residual of R remaining binary states. This insight can be derived by considering the process of uniquely assigning decimal digits. On the first level, there are nine decimal values, 1 to 9 . On the second level, each of these values is extended by adding digits 1 to 9 , yielding the combinations 1 . 1 to 1 . 9 , 2 . 1 to 2 . 9 , and so forth. In total, this results in 9·9=9·100·9=81 distinct states. The value 1 . 0 is not included explicitly, as this would constitute a redundant representation. At the third level, however, a peculiarity arises. Although 1 . 0 was not explicitly introduced previously, the states 1 . 01 to 1 . 09 must now be represented. Consequently, we append nine digits to each of the ten states from the preceding level, yielding a total of 9·10 ·9=810 states. In general, this relationship may be expressed as 9·10k·(10 −1) , indicating that one state in the final level is not explicitly stored. As an illustrative example, let a=7 , so that 2a=128 . It then follows that K=0 and R=38 . Hence, we can represent K+2=2 full decimal digits: 1 . 0 , 1 . 1 , ... , 2 . 0 , ... , 2 . 9 , ... , 9 . 0 , ... , 9 . 9 . The remaining 38 binary states are filled as follows. Since 38 =4·9+2 , four of the subsequent decimal digit sequences can be completed. For D=0 , filling proceeds from the upper end, yielding 9 . 91 to 9 . 99 , 9 . 81 to 9 . 89 , 9 . 71 to 9 . 79 , and 9 . 61 to 9 . 69 . Conversely, for D=1 , filling proceeds from the lower end, resulting in 1 . 01 to 1 . 09 , 1 . 11 to 1 . 19 , 1 . 21 to 1 . 29 , and 1 . 31 to 1 . 39 . The remaining two states are then used to partially fill the next group, namely 9 . 58 to 9 . 59 , and 1 . 41 to 1 . 42 , respectively. This procedure ensures a unique decimal representation. An efficient way to compute K and R is either to employ a lookup table for the five possible values of a , or to compute and threshold (2a−9)/81 . With the complete set of representable decimal values constructed, the integer value {︄uint(I)S=0 uint(︁I)︁S=1(7)
3 outlook 4 serves as an index into the sorted set, yielding the significand value d . The distinction by sign ensures the preservation of binary monotonicity. With all values in place, the final numerical value represented by the dekum is obtained as ⎧ ⎪ ⎪ ⎨ ⎪ ⎪ ⎩ {︄0S=0 NaR S=1(D,R,C,I) = 0n−1 (−1)S·d·10eotherwise. (8) Here, zero is represented as the all-zero bit string, while non-real numbers are expressed by the special symbol NaR (‘not a real’). 3 outlook This document serves solely to present the format. Further work is required to analyse its properties and refine the formalism, particularly with respect to the construction of the set of decimal values. It should be noted that not all properties characteristic of posits or takums can be retained; priority has been given to ensuring uniqueness, at the expense of straightforward convertibility. In its current form, the dekum format should already offer superior storage efficiency compared with existing decimal floating-point formats.