llama.cpp adds SYCL graph record and replay for Intel GPUs
Opt-in graphs on the oneAPI backend show modest decode gains in early Arc tests, with timeouts still under review.
By tensorOpt-in graphs on the oneAPI backend show modest decode gains in early Arc tests, with timeouts still under review.
By tensorTwo-node RPC tests show decode and prefill roughly 1.8x faster with lower intermediate memory use.
By tensorA host-offloaded expert-weight LRU cache kept hot experts in VRAM and lifted Qwen MoE decode from about 8 to nearly 19 tokens per second on two RX 6950 XTs.
By tensorA draft would let rights holders treat autonomous system use of assets differently from content a human user supplies at inference time.
By ttl