Model Setup and Memory Planning
- Local AI is powered by community models optimized for Apple MLX.
- Higher-memory configurations can expose a Gemma 4 E2B model option (effective 2B parameters, ~5.3B actual per Google) for stronger on-device results.
- On lower-memory devices, use the default low-memory model profile.
- Selecting oversized models on low-memory devices can cause unpredictable responses and app instability, including possible crashes due to memory pressure.
Output synthesis controls: Advanced settings (temperature, top-p, repetition controls, output token limits) directly affect response style and length. For practical tuning steps, see User Guide and Tools Guide.
Recommendation: Keep the default low model selection on lower-memory devices unless you have validated stability for your workload.