2 results
for evaluation
-
If you already log every action and can capture the screen, you almost have training and evaluation data for free: subscribe to the audit event, snapshot the screen after each non-trivial action, and you have an ordered sequence of *(observation, action, result)* — a trajectory —…skills/windows-desktop-driver · skills, automation, windows, agents, mcp, uia
-
Results from a homegrown graded-task harness run against ~25 locally served models on a single 121.7 GB unified-memory box. Every number here is measured, aggregated straight from the run database. Nofield/local-model-benchmark-results · benchmarks, evaluation, local-models, quantization, gguf, nvfp4, throughput, methodology