Layered local-LLM inference engine for agentic workloads: Ollama and llama.cpp behind one abstraction, context-memory (sink/window/evict + block retrieval), encrypted audit log. Native L3 serving layer in progress.

0 stars 3 forks 0 watchers Python Apache License 2.0
llama-cpp llm llm-inference local-llm
6 Open Issues Need Help Last updated: Aug 14, 2026

Open Issues Need Help

View All on GitHub
good first issue

Layered local-LLM inference engine for agentic workloads: Ollama and llama.cpp behind one abstraction, context-memory (sink/window/evict + block retrieval), encrypted audit log. Native L3 serving layer in progress.

Python
#llama-cpp#llm#llm-inference#local-llm

Layered local-LLM inference engine for agentic workloads: Ollama and llama.cpp behind one abstraction, context-memory (sink/window/evict + block retrieval), encrypted audit log. Native L3 serving layer in progress.

Python
#llama-cpp#llm#llm-inference#local-llm

Layered local-LLM inference engine for agentic workloads: Ollama and llama.cpp behind one abstraction, context-memory (sink/window/evict + block retrieval), encrypted audit log. Native L3 serving layer in progress.

Python
#llama-cpp#llm#llm-inference#local-llm
good first issue

Layered local-LLM inference engine for agentic workloads: Ollama and llama.cpp behind one abstraction, context-memory (sink/window/evict + block retrieval), encrypted audit log. Native L3 serving layer in progress.

Python
#llama-cpp#llm#llm-inference#local-llm

Layered local-LLM inference engine for agentic workloads: Ollama and llama.cpp behind one abstraction, context-memory (sink/window/evict + block retrieval), encrypted audit log. Native L3 serving layer in progress.

Python
#llama-cpp#llm#llm-inference#local-llm

Layered local-LLM inference engine for agentic workloads: Ollama and llama.cpp behind one abstraction, context-memory (sink/window/evict + block retrieval), encrypted audit log. Native L3 serving layer in progress.

Python
#llama-cpp#llm#llm-inference#local-llm