OpenAI ships GPT-5.6 as a three-model family: Luna, Terra, and Sol
GPT-5.6 hit general availability July 9 in three sizes, priced $1 to $5 per million input tokens, with a 1M context and tool calling that writes its own orchestration code.
⚡ Stream · page 2
Signal over noise: what actually matters in AI and 3D printing this week.
40 news items
GPT-5.6 hit general availability July 9 in three sizes, priced $1 to $5 per million input tokens, with a 1M context and tool calling that writes its own orchestration code.
OpenAI's GPT-Live replaces the ChatGPT voice model and can hand harder questions to GPT-5.5 in the background, keeping the conversation going while it waits for the answer.
Two weeks after lvl30 called Grok 4.5's private-beta claims unverifiable, there are checkable numbers: $2 per million input tokens, behind Fable 5 and GPT-5.5 on coding, far fewer tokens burned.
Build b9951 of llama.cpp drops an initial ExecuTorch backend, planting a flag for GGUF-format models on PyTorch's on-device inference stack and the edge hardware it targets.
Meta pulled Muse Image's Instagram feature after a privacy backlash, hours after a Reuters test found its detector missed 55% of cropped images.
Muse Spark 1.1 shipped July 9 alongside the first public Meta Model API, a US preview at $1.25/$4.25 per million tokens, putting Meta in the same paid-API business as OpenAI and Anthropic.
A community rig ran NVIDIA's Nemotron Puzzle 75B-A9B in NVFP4 across three power-capped RTX 3090s at 132 t/s decode, and asked why almost nobody ships models shaped for multi-24GB rigs.
OrcaSlicer 2.4.2 is a maintenance drop that patches the missing-preset warnings and cloud-sync regressions that stung users who upgraded past 2.4.1.
Unsloth published NVFP4 quantized checkpoints for Qwen3.6 27B and 35B-A3B that reportedly beat NVIDIA's reference quants on throughput, with FP8 KV cache support for longer contexts.
SK Hynix expects memory demand to outrun its supply beyond 2030, but the HBM crunch is not yet a reason to panic-buy a consumer GPU.
Z.ai's GLM-5.2 shipped under an MIT license and beats GPT-5.5 on coding benchmarks. It's also a 744B model that needs a server rack, not a desk.
Elon Musk says Grok 4.5 rivals Claude Opus. Nobody outside SpaceX and Tesla can test it, it's on no public benchmark, and there's no release date.
Mistral's CEO confirmed a new open-weight model with early access this month, as the company reportedly raises $3.5B on around $400M in annual revenue.
Ollama's latest build nearly doubles Gemma 4 token generation on Apple Silicon with multi-token prediction. It's automatic, needs no config, and doesn't change the output.
A wave of low-cost accelerator HATs and NPU-equipped boards is pushing real on-device inference down to single-board-computer prices. The interesting part isn't the TOPS number, it's where the work moves.
The site respawned, retooled for AI and 3D printing. What's launching, how often, and what the streams are all about.