Apple sues OpenAI, and the io deal is the crime scene
Apple's trade secret suit names OpenAI's hardware chief and claims candidates were told to bring actual Apple parts to job interviews. The AI phone might be built on Cupertino's homework.
⚡ Stream
Signal over noise: what actually matters in AI and 3D printing this week.
40 news items
Apple's trade secret suit names OpenAI's hardware chief and claims candidates were told to bring actual Apple parts to job interviews. The AI phone might be built on Cupertino's homework.
Moonshot AI's Kimi K3 shipped with open weights and enough capability to restart the DeepSeek argument: if a model you can download keeps pace with the frontier, what were the export controls buying?
llama.cpp build b10069 adds OpenCL broadcast support for Adreno MUL_MAT operations and fixes view offset handling for Adreno Q8_0 MUL_MAT in llama-server multi-stream mode.
A solo dev open-sourced NInfer, a from-scratch C++/CUDA engine that sustains 542 tok/s on Qwen3.6-35B-A3B over a full 65K-token decode, single request, on one RTX 5090.
Moonshot's Kimi K3 sits at #1 on the Code Arena Frontend leaderboard, but scores only ~39% on FrontierMath Tier 4 versus ~90% for top OpenAI and Anthropic models.
Running Qwen3.6 27B with Opencode? The model's appetite for context tokens can gut your usable window fast. Here's what's happening and how to push back.
From July 20, Claude Fable 5 is included in Max and Team Premium plans at 50% of limits. Pro and Team Standard stay on usage credits at $10/$50 per million tokens, softened by a one-time $100 credit.
Kimi K3 hit #1 on the Frontend Code Arena on July 16, beating Claude Fable 5. David Sacks calls it concerning, five weeks after backing the order that benched that same model.
Moonshot AI announced Kimi K3, a 2.8-trillion-parameter MoE model, on July 16. Open weights are promised by July 27. Until then it is API-only, and running it yourself will need serious cluster hardware.
llama.cpp build b10067 patches a bug where the DeepSeek-V4 routing table tensor was being incorrectly quantized, causing a type conversion failure that blocked anyone using llama-quantize on the model.
A Reddit user reportedly ran three concurrent Qwen3.5 122B sessions on a single Mac Studio, serving over 93% of prompt tokens from an on-disk KV cache instead of recomputing on GPU.
Google shipped a silent update to Gemma 4 fixing tool-calling bugs, truncated responses, and enabling Flash Attention 4 on Hopper GPUs for a reported 25–70% prefill speedup.
A Reddit report claims the 98GB DeepSeek-V4-Flash quant jumped from 2 to 7 t/s on a single 4060 Ti plus CPU, with no hardware changes between two llama.cpp builds.
One thousand Bondtech Founder's Edition INDX kits are now shipping for the Prusa CORE One. The standard Prusa Edition follows at end of July, with the full first batch out by end of August.
Build b9993 merges PR #25395, adding native hy_v3 architecture support so Tencent's Hunyuan 3 MoE runs locally in llama.cpp, with MTP speculative decoding included.
A llama.cpp fork from spiritbuun adds Variable Bit Rate KV cache quantization, dynamically degrading cache precision only when VRAM pressure demands it.
vLLM v0.25.0 ships MRv2 as the default execution path for all dense models and permanently removes the legacy PagedAttention implementation.
Colibrì is a pure-C, zero-dependency inference engine that runs GLM-5.2, a 744B MoE model, on a consumer machine with 25GB of RAM, by keeping dense weights in RAM and streaming routed experts from disk.
Joshua Bird's custom clay extruder uses an air compressor and stainless steel auger to print pottery. The process works, but shrinkage, abrasion, and a two-fire ceramic workflow make it genuinely hard.
Build b9963 lands proper multi-tile support for DeepSeek-OCR v1 with dynamic resolution, and unifies the image preprocessor for both v1 and v2. Multi-column documents now handled correctly.
AMD's 128GB unified-memory Ryzen AI Halo developer box launched at $3,999, about $700 under NVIDIA's Linux-only DGX Spark, sold in the US exclusively through Micro Center in-store pickup.
Claude's 5-hour and weekly limits got a clean-slate reset for all paid tiers this week, and included Fable 5 access was extended from July 7 to July 12 after subscriber backlash.
A day after GPT-5.6 went GA, OpenAI shipped ChatGPT Work: an in-ChatGPT agent that works across Slack, Teams, Drive, and email for hours at a stretch, metered like Codex.
Before the public launch, GPT-5.6 sat in a preview limited to about 20 government-vetted partners under the new US frontier-AI oversight process. The White House disputes calling the release an approval.