Skip to main content
Preference Modeling/HOT/

Getting 50 GB/S Back from the Apple Neural Engine

RTL performance erratum in the Apple M3 Neural Engine throttles DRAM weight streaming throughput down to 17–19 GB/s from the nominal 45–60 GB/s. Avoiding the problematic path in the kernel DMA engine's speculative prefetch ring increased Llama 3.2 1B token throughput from 10.0 to 24.3 tokens/s.

Current signal

Leaning Human

0 total evaluations · Watch the replay →

AI score

60

Human score

Human-AI gap

Preliminary signal. Fewer than five human judgments have been recorded, so this result may move substantially.

Your judgment is missing

Agree with the machine?

Vote across the five dimensions, see where you land against the crowd, and strengthen the public alignment dataset.

Cast your judgment