HW/FW security researcher & Demoscene elder.

I started having arguments online back on Fidonet and Usenet. I’m too tired to care now.

  • 15 Posts
  • 1.23K Comments
Joined 3 years ago
Aquileo | cake
Cake day: June 12th, 2023

Aquileo | help-circle




  • When I start llama-server I point it to the models-config that have unique max context sizes per model - and they’re allocated at their max size as soon as the server starts so since it comes up I will be able to use that context size too.

    I’m actually a bit unsure as to how you run it since you get OOMs during usage :)

    I also use the DCP plugin for Opencode to help manage the context cache and have less of a disruption as it gets compressed, but I wouldn’t need to for the above to work. When I hit the context limit the context would still get compressed.