akenji's lab
  • Blog
  • About
  • Categories
  • Tags
  • ja
akenji's lab

Elyza


Running Elyza models on GPU using llama-cpp-python

 Posted on May 3, 2024  |    • Other languages: ja

Motivation

Quantization is essential to run LLM on the local workstation (12-16 GB of GPU memory). In this post, I summarize my attempt to maximize GPU resources using llama-cpp-python.

The content includes some of my mistakes, as I got into some areas due to my lack of understanding.

[Read More]
Elyza  llama-cpp-python  GPU 

 • © 2026  •  akenji's lab

Hugo v0.147.1 powered  •  Theme Beautiful Hugo adapted from Beautiful Jekyll