Not every workflow can ship its data to a cloud API. Local model deployment puts a capable LLM on hardware the client owns — Gemma for general use, specialized models where they win, and a clear answer on what runs locally vs. what still needs the cloud.
Stack runs on Ollama and llama.cpp where appropriate, with thin PHP wrappers for the application layer.
Stack & Tools
Gemma
Ollama
llama.cpp
Linux