Repository navigation
llvm.fma.f16 intrinsic is expanded incorrectly on targets without native half FMA support #98389
Description
Activity
- addedfloating-pointFloating-point mathFloating-point mathllvm:SelectionDAGSelectionDAGISel as wellSelectionDAGISel as welland removed
on Aug 7, 2024 It may be preferable to lower to a
fmaf16libcall in cases where there is also no hardwaref64to avoid the 128 bit multiplication needed for soft floatf64 * f64.- added a commit that references this issue
on Dec 17, 2025 - added a commit that references this issue
on Dec 17, 2025 @llvm/issue-subscribers-backend-risc-v
Author: None (beetrees)
Details
Consider the following LLVM IR: ```llvm declare half @llvm.fma.f16(half %a, half %b, half %c)define half @do_fma(half %a, half %b, half %c) {
%res = call half @llvm.fma.f16(half %a, half %b, half %c)
ret half %res
}On targets without native `half` FMA support, LLVM turns this into the equivalent of: ```llvm declare float @<!-- -->llvm.fma.f32(float %a, float %b, float %c) define half @<!-- -->do_fma(half %a, half %b, half %c) { %a_f32 = fpext half %a to float %b_f32 = fpext half %b to float %c_f32 = fpext half %c to float %res_f32 = call float @<!-- -->llvm.fma.f32(float %a_f32, float %b_f32, float %c_f32) %res = fptrunc float %res_f32 to half ret half %res }This is a miscompilation, however, as
floatdoes not have enough precision to do a fused-multiply-add forhalfwithout double rounding becoming an issue. For instance (raw bits of eachhalfare in brackets):do_fma(48.34375 (0x520b), 0.000013887882 (0x00e9), 0.12438965 (0x2ff6)) = 0.12512207 (0x3001), but LLVM's lowering tofloatFMA gives an incorrect result of0.125 (0x3000).A correct lowering would need to use
double(or larger): adoubleFMA is not required asdoubleis large enough to represent the result ofhalf * halfwithout any rounding. In summary, a correct lowering would look something like this:declare double @<!-- -->llvm.fmuladd.f64(double %a, double %b, double %c) define half @<!-- -->do_fma(half %a, half %b, half %c) { %a_f64 = fpext half %a to double %b_f64 = fpext half %b to double %c_f64 = fpext half %c to double %res_f64 = call double @<!-- -->llvm.fmuladd.f64(double %a_f64, double %b_f64, double %c_f64) %res = fptrunc double %res_f64 to half ret half %res }
The same fix will need to be applied for GlobalIsel
- added a commit that references this issue
on Dec 19, 2025 - marked [clang][X86] Wrong result for __builtin_elementwise_fma on _Float16 #128450 as a duplicate of this issue
on Jan 28, 2026 - added a commit that references this issue
on Jun 14, 2026 - added 5 commits that reference this issue
on Aug 25, 2026 - added a commit that references this issue
on Aug 25, 2026 - added a commit that references this issue
on Aug 26, 2026 - added a commit that references this issue
on Sep 8, 2026
Consider the following LLVM IR:
On targets without native
halfFMA support, LLVM turns this into the equivalent of:This is a miscompilation, however, as
floatdoes not have enough precision to do a fused-multiply-add forhalfwithout double rounding becoming an issue. For instance (raw bits of eachhalfare in brackets):do_fma(48.34375 (0x520b), 0.000013887882 (0x00e9), 0.12438965 (0x2ff6)) = 0.12512207 (0x3001), but LLVM's lowering tofloatFMA gives an incorrect result of0.125 (0x3000).A correct lowering would need to use
double(or larger): adoubleFMA is not required asdoubleis large enough to represent the result ofhalf * halfwithout any rounding. In summary, a correct lowering would look something like this: