LM: Difference between revisions

From Fundamental Ramen
Jump to navigation Jump to search
 
(157 intermediate revisions by the same user not shown)
Line 1: Line 1:
= Build environment =
= Topics =
 
== hf (model management) ==
* https://huggingface.co/docs/huggingface_hub/package_reference/environment_variables


{| class="wikitable"
{| class="wikitable"
! Purpose || Command
! 🤖 Agent  || 🧠 Model  || 🚀 Engine  || ⚙️ GPU & Driver || Evaluation
|-
|-
| cache management ||
| style="width:200px;vertical-align:top" |
<syntaxhighlight lang="bash">
: [[LM/ZooCode|ZooCode]]
hf cache list
: [[LM/jrswab_axe|axe]]
hf cache rm <model id>
: [[LM/MCP|MCP]]
hf cache prune
| style="width:200px;vertical-align:top" |
</syntaxhighlight>
: [[LM/hf|hf (Model Management)]]
|-
: [[LM/Qwen 3.8 27B|Qwen 3.8 27B]]
| fix WiFi problem ||
: [[LM/Gemma 4|Gemma 4]]
<syntaxhighlight lang="bash">
: [[LM/Devstral Small 2 24B|Devstral Small 2 24B]]
HF_XET_FIXED_DOWNLOAD_CONCURRENCY=10 hf download "unsloth/Qwen3-Coder-30B-A3B-Instruct-GGUF" --include "*UD-Q4_K_XL*"
: [[LM/Experience|Experience]]
HF_XET_FIXED_DOWNLOAD_CONCURRENCY=10 hf download "unsloth/Devstral-Small-2-24B-Instruct-2512-GGUF" --include "*UD-Q4_K_XL*"
| style="width:200px;vertical-align:top" |
</syntaxhighlight>
: [[LM/llama-swap|llama-swap]]
|-
: [[LM/vLLM|vLLM]]
| search ||
: [[LM/ComfyUI|ComfyUI]]
<syntaxhighlight lang="bash">
: [[LM/llama.cpp for SYCL|llama.cpp for SYCL]]
hf models ls --search "gemma-4" --apps llama.cpp --sort downloads --limit 10
: [[LM/llama.cpp for CUDA|llama.cpp for CUDA]]
hf models ls --search "coder" --apps llama.cpp --sort downloads --limit 10
: [[LM/llama.cpp for OpenVINO|llama.cpp for OpenVINO]]
hf models ls --search "unsloth" --apps llama.cpp --sort downloads --limit 10
: [[LM/llama.cpp for TurboQuant|llama.cpp for TurboQuant]]
</syntaxhighlight>
: [[LM/SGLang|SGLang]]
|-
| style="width:200px;vertical-align:top" |
| optimize Qwen3-Coder ||
: Intel Arc Pro B70
<syntaxhighlight lang="bash">
:: [[LM/Install driver|Install driver]]
hf models ls --search "coder" --apps llama.cpp --sort downloads --limit 1 --format json | jq .
:: [[LM/Install oneAPI|Install oneAPI (SYCL)]]
hf models ls -h "unsloth/Qwen3-Coder-30B-A3B-Instruct-GGUF"
:: [[LM/Install OpenVINO|Install OpenVINO]]
hf download "unsloth/Qwen3-Coder-30B-A3B-Instruct-GGUF" --include "*UD-Q4_K_XL*"
:: [[LM/.bashrc for B70|.bashrc for B70]]
</syntaxhighlight>
:: [[LM/Intel Pro Arc B70|Experiment]]
|-
: NVIDIA RTX 3060
| optimize gemma-4-E4B ||
:: [[LM/Install NVIDIA driver|Install driver]]
<syntaxhighlight lang="bash">
: AMD Ryzen AI 7 350
hf models ls -h unsloth/gemma-4-E4B-it-qat-GGUF
:: [[LM/Ryzen AI 7 350|Experiment]]
hf download unsloth/gemma-4-E4B-it-qat-GGUF --include "*UD-Q4_K_XL*"
: Tools
hf download unsloth/gemma-4-E4B-it-qat-GGUF --include "mmproj-BF16.gguf"
:: [[LM/Install nvtop|Install nvtop]]
hf download unsloth/gemma-4-E4B-it-qat-GGUF --include "mtp-gemma-4-E4B-it.gguf"
:: [[LM/lact|lact]]
</syntaxhighlight>
:: [[LM/PCIe Trouble Shooting|PCIe Trouble Shooting]]
|-
: Comparison
| optimize gemma-4-12B ||
:: [[LM/GPU Comparison]]
<syntaxhighlight lang="bash">
| style="width:200px;vertical-align:top" |
hf models ls -h unsloth/gemma-4-12B-it-qat-GGUF
: [[LM/Performance|Performance]]
hf download unsloth/gemma-4-12B-it-qat-GGUF --include "*UD-Q4_K_XL*"
hf download unsloth/gemma-4-12B-it-qat-GGUF --include "mmproj-BF16.gguf"
hf download unsloth/gemma-4-12B-it-qat-GGUF --include "mtp-gemma-4-12B-it.gguf"
</syntaxhighlight>
|-
| optimize Muse Glimmer ||
<syntaxhighlight lang="bash">
hf models ls -h unsloth/Muse-Glimmer-30B-GGUF
hf download unsloth/Muse-Glimmer-30B-GGUF --include "*UD-Q4_K_XL*"
hf download unsloth/Muse-Glimmer-30B-GGUF --include "*UD-Q5_K_XL*"
hf download unsloth/Muse-Glimmer-30B-GGUF --include "*UD-Q6_K_XL*"
hf download unsloth/Muse-Glimmer-30B-GGUF --include "dflash-kquant.gguf"
hf download unsloth/Muse-Glimmer-30B-GGUF --include "mmproj-kquant.gguf"
</syntaxhighlight>
|-
| ||
<syntaxhighlight lang="bash">
tree ~/.cache/huggingface/hub/models--unsloth--Devstral-Small-2-24B-Instruct-2512-GGUF/snapshots
tree ~/.cache/huggingface/hub/models--unsloth--Qwen3-Coder-30B-A3B-Instruct-GGUF/snapshots
</syntaxhighlight>
|}
|}


== Ubuntu ==
= News =
 
* [https://huggingface.co/datasets/ScaleAI/SWE-bench_Pro?leaderboard_max_params=32B SWE Pro Benchmark for models < 32B]
* [https://storage.openvinotoolkit.org/repositories/openvino/packages/2026.3/linux/ OpenVINO 2026.3 supports Ubuntu 26] - 2026-08-04
 
= Models management =
 
<quickmmd name="models-handling-infra">
flowchart LR
  A(LiteLLM)
  B1(llama-swap)
  B2(vLLM<br>Qwen 3.8 27B)
  C(llama.cpp<br>Gemma 4 E4B)
 
  A --> B1 & B2
  B1 --> C
</quickmmd>
 
* llama-swap can launch model, cannot manage users.
* LiteLLM cannot launch model, can manage users.
 
= Decision Tree for tinkering with Intel Arc Pro B70 =
 
<quickmmd name="tinkering-b70">
flowchart TD
  subgraph Hardware
    A1(Intel Arc Pro B70):::node-okay
    A2(Asus ProArt B850 WiFi NEO):::node-okay
    A3(AMD Ryzen 5 9600X):::node-okay
  end
 
  subgraph OS
    B1(Install<br>Ubuntu 26.04):::node-okay
    A1 & A2 & A3 --> B1
  end
 
  subgraph Driver
    C1(Install PPA driver<br>26.27.39122.14):::node-okay
    C2(Install oneAPI<br>2026.1.1.20260724):::node-okay
    C3(Install OpenVINO<br>2026.3.0<br>hard to install):::node-warn
    C1 --> C2 & C3
    B1 --> C1
  end
 
  subgraph Runtime
    D1[Build llama.cpp for<br>SYCL]:::node-okay
    D2[Build llama.cpp for<br>OpenVINO<br>not stable yet]:::node-warn
    %% D3[Build SGLang for<br>SYCL]:::node-todo
    %% D4[Build SGLang for<br>OpenVINO]:::node-todo
    C2 --> D1
    C3 --> D2
    %% C2 --> D3
    %% C3 --> D4
  end


* [[LLM/Intel Pro Arc B70]]
  subgraph Tools
* [[LLM/Ryzen AI 7 350]] (XDNA)
    T1(Install<br>pipx)
    T2(Install<br>hf)
    T3(Install<br>nvtop)
    T4(Install<br>tmux)
    T5(Install<br>lact)
    T6(Install<br>llama-swap)
    T1 --> T2
    B1 --> T1
    B1 --> T3
    B1 --> T4
    B1 --> T5
    T5 --> T6
  end


== Agents ==
  subgraph Models
    M1(Qwen3.8-27B)
    M2(Devstral-Small-2-24B)
    T2 --> M1 & M2
  end


* [[LLM/jrswab_axe]]
  subgraph config
    V1(Optimize<br>.bashrc)
    V2(Optimize<br>lmsw.yaml)
    C2 --> V1
    C3 --> V1
    T2 --> V1
    T4 ---> V1
    T6 ---> V2
  end


== Wishlist ==
  subgraph Service
    Z(service<br>llama-swap \<br>  -config lmsw.yaml \<br>  -listen 0.0.0.0:9876)
    M1 --> Z
    D1 --> Z
    V2 --> Z
  end


* [[LLM/prompt for coding]]
  classDef node-okay fill:#efe,stroke:#393
  classDef node-warn fill:#fff0e0,stroke:#d63
  classDef node-fail fill:#fee,stroke:#d33
  classDef node-todo fill:#eee,stroke:#777
</quickmmd>

Latest revision as of 09:49, 21 September 2026

Topics

🤖 Agent  🧠 Model  🚀 Engine  ⚙️ GPU & Driver  Evaluation
ZooCode
axe
MCP
hf (Model Management)
Qwen 3.8 27B
Gemma 4
Devstral Small 2 24B
Experience
llama-swap
vLLM
ComfyUI
llama.cpp for SYCL
llama.cpp for CUDA
llama.cpp for OpenVINO
llama.cpp for TurboQuant
SGLang
Intel Arc Pro B70
Install driver
Install oneAPI (SYCL)
Install OpenVINO
.bashrc for B70
Experiment
NVIDIA RTX 3060
Install driver
AMD Ryzen AI 7 350
Experiment
Tools
Install nvtop
lact
PCIe Trouble Shooting
Comparison
LM/GPU Comparison
Performance

News

Models management

  • llama-swap can launch model, cannot manage users.
  • LiteLLM cannot launch model, can manage users.

Decision Tree for tinkering with Intel Arc Pro B70