v11: 方向修正(hot expert 實體放 RAM,非 page 轉換)+ b9acf138 vanilla 基線 82.1 tok/s / peak RSS 2.674 GiB a035243
Auto Upload Agent commited on
How to use HelloSun/SmallThinker4b with llama-cpp-python:
# !pip install llama-cpp-python
from llama_cpp import Llama
llm = Llama.from_pretrained(
repo_id="HelloSun/SmallThinker4b",
filename="{{GGUF_FILE}}",
)
output = llm( "Once upon a time,", max_tokens=512, echo=True ) print(output)