Skip to content

measure(MODEL-NEMOTRON-H-ABI-A2P): the first Nemotron ratios -- decode 0.001392x, load 2.12x faster, and the GPU idle for 93.7% of our decode (#1250) - #1251

Open
localai-bot wants to merge 29 commits into
mainfrom
row/MODEL-NEMOTRON-H-ABI-A2P-speed
Open

measure(MODEL-NEMOTRON-H-ABI-A2P): the first Nemotron ratios -- decode 0.001392x, load 2.12x faster, and the GPU idle for 93.7% of our decode (#1250)#1251
localai-bot wants to merge 29 commits into
mainfrom
row/MODEL-NEMOTRON-H-ABI-A2P-speed