LM: Difference between revisions

From Fundamental Ramen
Jump to navigation Jump to search
 
(152 intermediate revisions by the same user not shown)
Line 1: Line 1:
= Topics =
{| class="wikitable"
! 🤖 Agent  || 🧠 Model  || 🚀 Engine  || ⚙️ GPU & Driver || Evaluation
|-
| style="width:200px;vertical-align:top" |
: [[LM/ZooCode|ZooCode]]
: [[LM/jrswab_axe|axe]]
: [[LM/MCP|MCP]]
| style="width:200px;vertical-align:top" |
: [[LM/hf|hf (Model Management)]]
: [[LM/Qwen 3.8 27B|Qwen 3.8 27B]]
: [[LM/Gemma 4|Gemma 4]]
: [[LM/Devstral Small 2 24B|Devstral Small 2 24B]]
: [[LM/Experience|Experience]]
| style="width:200px;vertical-align:top" |
: [[LM/llama-swap|llama-swap]]
: [[LM/vLLM|vLLM]]
: [[LM/ComfyUI|ComfyUI]]
: [[LM/llama.cpp for SYCL|llama.cpp for SYCL]]
: [[LM/llama.cpp for CUDA|llama.cpp for CUDA]]
: [[LM/llama.cpp for OpenVINO|llama.cpp for OpenVINO]]
: [[LM/llama.cpp for TurboQuant|llama.cpp for TurboQuant]]
: [[LM/SGLang|SGLang]]
| style="width:200px;vertical-align:top" |
: Intel Arc Pro B70
:: [[LM/Install driver|Install driver]]
:: [[LM/Install oneAPI|Install oneAPI (SYCL)]]
:: [[LM/Install OpenVINO|Install OpenVINO]]
:: [[LM/.bashrc for B70|.bashrc for B70]]
:: [[LM/Intel Pro Arc B70|Experiment]]
: NVIDIA RTX 3060
:: [[LM/Install NVIDIA driver|Install driver]]
: AMD Ryzen AI 7 350
:: [[LM/Ryzen AI 7 350|Experiment]]
: Tools
:: [[LM/Install nvtop|Install nvtop]]
:: [[LM/lact|lact]]
:: [[LM/PCIe Trouble Shooting|PCIe Trouble Shooting]]
: Comparison
:: [[LM/GPU Comparison]]
| style="width:200px;vertical-align:top" |
: [[LM/Performance|Performance]]
|}
= News =
* [https://huggingface.co/datasets/ScaleAI/SWE-bench_Pro?leaderboard_max_params=32B SWE Pro Benchmark for models < 32B]
* [https://storage.openvinotoolkit.org/repositories/openvino/packages/2026.3/linux/ OpenVINO 2026.3 supports Ubuntu 26] - 2026-08-04
= Models management =
<quickmmd name="models-handling-infra">
flowchart LR
  A(LiteLLM)
  B1(llama-swap)
  B2(vLLM<br>Qwen 3.8 27B)
  C(llama.cpp<br>Gemma 4 E4B)
  A --> B1 & B2
  B1 --> C
</quickmmd>
* llama-swap can launch model, cannot manage users.
* LiteLLM cannot launch model, can manage users.
= Decision Tree for tinkering with Intel Arc Pro B70 =
= Decision Tree for tinkering with Intel Arc Pro B70 =


<quickmmd name="tinkering-b70">
<quickmmd name="tinkering-b70">
flowchart TB
flowchart TD
   A[Got B70]
   subgraph Hardware
  B1[Install Ubuntu 24.04]
    A1(Intel Arc Pro B70):::node-okay
   B2[Install Ubuntu 26.04]
    A2(Asus ProArt B850 WiFi NEO):::node-okay
</quickmmd>
    A3(AMD Ryzen 5 9600X):::node-okay
   end


= Build environment =
  subgraph OS
    B1(Install<br>Ubuntu 26.04):::node-okay
    A1 & A2 & A3 --> B1
  end


== hf (model management) ==
  subgraph Driver
* https://huggingface.co/docs/huggingface_hub/package_reference/environment_variables
    C1(Install PPA driver<br>26.27.39122.14):::node-okay
* https://www.datalearner.com/en/leaderboards/category/code?benchmark=SWE-bench+Verified&modelSize=34b&licenseType=open
    C2(Install oneAPI<br>2026.1.1.20260724):::node-okay
* https://huggingface.co/datasets/ScaleAI/SWE-bench_Pro
    C3(Install OpenVINO<br>2026.3.0<br>hard to install):::node-warn
    C1 --> C2 & C3
    B1 --> C1
  end


{| class="wikitable"
  subgraph Runtime
! Purpose || Command
    D1[Build llama.cpp for<br>SYCL]:::node-okay
|-
    D2[Build llama.cpp for<br>OpenVINO<br>not stable yet]:::node-warn
| cache management ||
    %% D3[Build SGLang for<br>SYCL]:::node-todo
<syntaxhighlight lang="bash">
    %% D4[Build SGLang for<br>OpenVINO]:::node-todo
hf cache list
    C2 --> D1
hf cache rm <model id>
    C3 --> D2
hf cache prune
    %% C2 --> D3
</syntaxhighlight>
    %% C3 --> D4
|-
  end
| fix WiFi problem ||
<syntaxhighlight lang="bash">
HF_XET_FIXED_DOWNLOAD_CONCURRENCY=10 hf download "unsloth/Qwen3-Coder-30B-A3B-Instruct-GGUF" --include "*UD-Q4_K_XL*"
HF_XET_FIXED_DOWNLOAD_CONCURRENCY=10 hf download "unsloth/Devstral-Small-2-24B-Instruct-2512-GGUF" --include "*UD-Q4_K_XL*"
</syntaxhighlight>
|-
| search ||
<syntaxhighlight lang="bash">
hf models ls --search "gemma-4" --apps llama.cpp --sort downloads --limit 10
hf models ls --search "coder" --apps llama.cpp --sort downloads --limit 10
hf models ls --search "unsloth" --apps llama.cpp --sort downloads --limit 10
</syntaxhighlight>
|-
| optimize Qwen3-Coder ||
<syntaxhighlight lang="bash">
hf models ls --search "coder" --apps llama.cpp --sort downloads --limit 1 --format json | jq .
hf models ls -h "unsloth/Qwen3-Coder-30B-A3B-Instruct-GGUF"
hf download "unsloth/Qwen3-Coder-30B-A3B-Instruct-GGUF" --include "*UD-Q4_K_XL*"
</syntaxhighlight>
|-
| optimize gemma-4-E4B ||
<syntaxhighlight lang="bash">
hf models ls -h unsloth/gemma-4-E4B-it-qat-GGUF
hf download unsloth/gemma-4-E4B-it-qat-GGUF --include "*UD-Q4_K_XL*"
hf download unsloth/gemma-4-E4B-it-qat-GGUF --include "mmproj-BF16.gguf"
hf download unsloth/gemma-4-E4B-it-qat-GGUF --include "mtp-gemma-4-E4B-it.gguf"
</syntaxhighlight>
|-
| optimize gemma-4-12B ||
<syntaxhighlight lang="bash">
hf models ls -h unsloth/gemma-4-12B-it-qat-GGUF
hf download unsloth/gemma-4-12B-it-qat-GGUF --include "*UD-Q4_K_XL*"
hf download unsloth/gemma-4-12B-it-qat-GGUF --include "mmproj-BF16.gguf"
hf download unsloth/gemma-4-12B-it-qat-GGUF --include "mtp-gemma-4-12B-it.gguf"
</syntaxhighlight>
|-
| optimize Muse Glimmer ||
<syntaxhighlight lang="bash">
hf models ls -h unsloth/Muse-Glimmer-30B-GGUF
hf download unsloth/Muse-Glimmer-30B-GGUF --include "*UD-Q4_K_XL*"
hf download unsloth/Muse-Glimmer-30B-GGUF --include "*UD-Q5_K_XL*"
hf download unsloth/Muse-Glimmer-30B-GGUF --include "*UD-Q6_K_XL*"
hf download unsloth/Muse-Glimmer-30B-GGUF --include "dflash-kquant.gguf"
hf download unsloth/Muse-Glimmer-30B-GGUF --include "mmproj-kquant.gguf"
</syntaxhighlight>
|-
|  ||
<syntaxhighlight lang="bash">
tree ~/.cache/huggingface/hub/models--unsloth--Devstral-Small-2-24B-Instruct-2512-GGUF/snapshots
tree ~/.cache/huggingface/hub/models--unsloth--Qwen3-Coder-30B-A3B-Instruct-GGUF/snapshots
</syntaxhighlight>
|}


== Ubuntu ==
  subgraph Tools
    T1(Install<br>pipx)
    T2(Install<br>hf)
    T3(Install<br>nvtop)
    T4(Install<br>tmux)
    T5(Install<br>lact)
    T6(Install<br>llama-swap)
    T1 --> T2
    B1 --> T1
    B1 --> T3
    B1 --> T4
    B1 --> T5
    T5 --> T6
  end


* [[LLM/Intel Pro Arc B70]]
  subgraph Models
* [[LLM/Ryzen AI 7 350]] (XDNA)
    M1(Qwen3.8-27B)
    M2(Devstral-Small-2-24B)
    T2 --> M1 & M2
  end


== Agents ==
  subgraph config
    V1(Optimize<br>.bashrc)
    V2(Optimize<br>lmsw.yaml)
    C2 --> V1
    C3 --> V1
    T2 --> V1
    T4 ---> V1
    T6 ---> V2
  end


* [[LLM/jrswab_axe]]
  subgraph Service
    Z(service<br>llama-swap \<br>  -config lmsw.yaml \<br>  -listen 0.0.0.0:9876)
    M1 --> Z
    D1 --> Z
    V2 --> Z
  end


== Wishlist ==
  classDef node-okay fill:#efe,stroke:#393
 
  classDef node-warn fill:#fff0e0,stroke:#d63
* [[LLM/prompt for coding]]
  classDef node-fail fill:#fee,stroke:#d33
  classDef node-todo fill:#eee,stroke:#777
</quickmmd>

Latest revision as of 09:49, 21 September 2026

Topics

🤖 Agent  🧠 Model  🚀 Engine  ⚙️ GPU & Driver  Evaluation
ZooCode
axe
MCP
hf (Model Management)
Qwen 3.8 27B
Gemma 4
Devstral Small 2 24B
Experience
llama-swap
vLLM
ComfyUI
llama.cpp for SYCL
llama.cpp for CUDA
llama.cpp for OpenVINO
llama.cpp for TurboQuant
SGLang
Intel Arc Pro B70
Install driver
Install oneAPI (SYCL)
Install OpenVINO
.bashrc for B70
Experiment
NVIDIA RTX 3060
Install driver
AMD Ryzen AI 7 350
Experiment
Tools
Install nvtop
lact
PCIe Trouble Shooting
Comparison
LM/GPU Comparison
Performance

News

Models management

  • llama-swap can launch model, cannot manage users.
  • LiteLLM cannot launch model, can manage users.

Decision Tree for tinkering with Intel Arc Pro B70