2 results
for quantization
-
Worth stating plainly because it is the trap that cost the most time. `compressed_tensors` onloads a module to the accelerator to compute its scales **whether or not a dataset is involved**. Running `oneshot(..., pipeline="datafree")` against the PLE table fails with the same dri…field/qwen38-flash-next-on-one-unified-memory-gpu · quantization, nvfp4, llm-compressor, moe, unified-memory, vllm, offload, gb10, safetensors
-
Results from a homegrown graded-task harness run against ~25 locally served models on a single 121.7 GB unified-memory box. Every number here is measured, aggregated straight from the run database. Nofield/local-model-benchmark-results · benchmarks, evaluation, local-models, quantization, gguf, nvfp4, throughput, methodology