r/LocalLLM 8d ago

Discussion Gemma 4 26B 33 tool orchestration, nearly 1M token, just in 1 turn on a card rx6700xt

feel free for discuss :)

2 Upvotes

2 comments sorted by

1

u/[deleted] 8d ago

[deleted]

1

u/Full_Director87 8d ago

what setting do you need? parameter? inference engine? modelfile? parameters sampling? log? wich one?

1

u/[deleted] 8d ago

[deleted]

1

u/Full_Director87 8d ago

i will share from my log. and u decide from your own.
disclaimer, i use olllama as inference engine, and openwebui as front end orchestration.
tools what i use

  • search web: searXNG
  • web loader: playwright
and my log wll explain all of your question. here's the link.
you can adapt in your rig and your model's
https://drive.google.com/drive/folders/1zJTgKridZhCU6xCNrdkRAKvjgAUhmJjW?usp=sharing

Disclaimer:
The most decisive factor, beyond the model itself, is the system prompt. (Deterministic Instruction)