<?xml version="1.0" encoding="utf-8" standalone="yes"?><rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom"><channel><title>moyutianzun's blog</title><link>https://moyutianzun.cn/en/</link><description>Recent content on moyutianzun's blog</description><generator>Hugo</generator><language>en</language><copyright>moyutianzun</copyright><lastBuildDate>Mon, 24 Aug 2026 09:00:00 +0800</lastBuildDate><atom:link href="https://moyutianzun.cn/en/index.xml" rel="self" type="application/rss+xml"/><item><title>moyutianzun.com</title><link>https://moyutianzun.cn/en/projects/moyutianzun-com/</link><pubDate>Mon, 24 Aug 2026 00:00:00 +0800</pubDate><guid>https://moyutianzun.cn/en/projects/moyutianzun-com/</guid><description>This very site: Hugo + PaperMod, 64 posts migrated losslessly from a Halo backup, four sections (news/blog/projects/about), llms.txt open for AI retrieval.</description></item><item><title>Loop 2026-08-24: this column starts</title><link>https://moyutianzun.cn/en/loop/2026-08-24/</link><pubDate>Mon, 24 Aug 2026 09:00:00 +0800</pubDate><guid>https://moyutianzun.cn/en/loop/2026-08-24/</guid><description>The Loop column is live. Everything the loop system produces lands here — maybe a news digest today, a cross-topic analysis tomorrow; the format follows wherever the loop evolves.</description></item><item><title>About</title><link>https://moyutianzun.cn/en/about/</link><pubDate>Mon, 24 Aug 2026 00:00:00 +0800</pubDate><guid>https://moyutianzun.cn/en/about/</guid><description>&lt;section class="mast about-mast"&gt;&#10; &lt;h1&gt;&lt;span class="mk"&gt;think&lt;/span&gt;About&lt;/h1&gt;&#10; &lt;p class="intro"&gt;I'm &lt;strong&gt;moyutianzun&lt;/strong&gt;. Most of what I care about converges on one question: &lt;strong&gt;how can context accumulate and compound&lt;/strong&gt;. I explore new paradigms of &lt;strong&gt;RSI&lt;/strong&gt; and &lt;strong&gt;context engineering&lt;/strong&gt; toward that question — each card below is a direction. Directions are broad and will evolve, but they all face the same question.&lt;/p&gt;&#10; &lt;div class="facts"&gt;&#10; &lt;div class="fact"&gt;&lt;span class="fk"&gt;Pacific time&lt;/span&gt;&lt;span class="fv" data-clock&gt;—&lt;/span&gt;&lt;/div&gt;&#10; &lt;div class="fact"&gt;&lt;span class="fk"&gt;Now&lt;/span&gt;&lt;span class="fv tmp" data-temp&gt;—&lt;/span&gt;&lt;/div&gt;&#10; &lt;div class="fact"&gt;&lt;span class="fk"&gt;Directions&lt;/span&gt;&lt;span class="fv"&gt;4&lt;/span&gt;&lt;/div&gt;&#10; &lt;div class="fact"&gt;&lt;span class="fk"&gt;Projects&lt;/span&gt;&lt;span class="fv"&gt;3&lt;/span&gt;&lt;/div&gt;&#10; &lt;/div&gt;&#10;&lt;/section&gt;&#10;&#10;&lt;section class="sec" id="directions"&gt;&#10; &lt;div class="sh"&gt;&lt;b&gt;Directions&lt;/b&gt;&lt;/div&gt;&#10; &lt;article class="dir" id="d01"&gt;&#10; &lt;header&gt;&lt;span class="no"&gt;01&lt;/span&gt;&lt;h3&gt;Funny Eval&lt;/h3&gt;&lt;span class="ds"&gt;in progress&lt;/span&gt;&lt;/header&gt;&lt;p class="lead"&gt;Leaderboards don't show what a model can actually do. I want to force the gap into the open &lt;strong&gt;visually and playfully&lt;/strong&gt; — a task is worth building only if you can see who won at a glance.&lt;/p&gt;</description></item><item><title>垂类系统的agent如何设计</title><link>https://moyutianzun.cn/en/blog/chui-lei-xi-tong-de-agentru-he-she-ji/</link><pubDate>Mon, 29 Jun 2026 07:59:28 +0800</pubDate><guid>https://moyutianzun.cn/en/blog/chui-lei-xi-tong-de-agentru-he-she-ji/</guid><description>&lt;h2 id="核心矛盾1垂直深度-vs-水平覆盖"&gt;核心矛盾1：垂直深度 vs 水平覆盖&lt;a class="hl-a" href="#%e6%a0%b8%e5%bf%83%e7%9f%9b%e7%9b%be1%e5%9e%82%e7%9b%b4%e6%b7%b1%e5%ba%a6-vs-%e6%b0%b4%e5%b9%b3%e8%a6%86%e7%9b%96" aria-hidden="true"&gt;#&lt;/a&gt;&lt;/h2&gt;&#10;&lt;p&gt;一个「霸总甜宠」skill 要写到能指导 LLM 生成高质量脚本，需要非常具体的规则——比如「第一集前两分钟必须完成阶级落差建立」「打脸节奏三集一小五集一大」。但这些规则对一个 30 秒广告完全没意义，对 80 集长剧又太稀疏。&lt;/p&gt;</description></item><item><title>每日github项目解析：（二）20260605 github robot和agent soul</title><link>https://moyutianzun.cn/en/blog/mei-ri-githubxiang-mu-jie-xi-er-20260605/</link><pubDate>Fri, 05 Jun 2026 09:47:40 +0800</pubDate><guid>https://moyutianzun.cn/en/blog/mei-ri-githubxiang-mu-jie-xi-er-20260605/</guid><description>&lt;h1 id="github-robot"&gt;github robot&lt;a class="hl-a" href="#github-robot" aria-hidden="true"&gt;#&lt;/a&gt;&lt;/h1&gt;&#10;&lt;p&gt;&lt;a href="https://github.com/pbakaus/agent-reviews"&gt;https://github.com/pbakaus/agent-reviews&lt;/a&gt;&lt;/p&gt;&#10;&lt;p&gt;&lt;strong&gt;这个项目让我想到一个有意思的场景：一个让人血压升高的下午&lt;/strong&gt;&lt;/p&gt;&#10;&lt;p&gt;你提了一个 PR，信心满满。几秒钟后，Copilot 来了，CodeRabbit 来了，Cursor Bugbot 也来了。它们在你的代码行上密密麻麻留下几十条评论：这里可能空指针，那里命名不规范，这个函数复杂度超标。你认认真真改了一轮，git push。&lt;/p&gt;</description></item><item><title>每日github项目解析：（一）20260604</title><link>https://moyutianzun.cn/en/blog/mei-ri-githubxiang-mu-jie-xi-yi-20260604/</link><pubDate>Thu, 04 Jun 2026 10:00:30 +0800</pubDate><guid>https://moyutianzun.cn/en/blog/mei-ri-githubxiang-mu-jie-xi-yi-20260604/</guid><description>&lt;h1 id="headroom"&gt;headroom&lt;a class="hl-a" href="#headroom" aria-hidden="true"&gt;#&lt;/a&gt;&lt;/h1&gt;&#10;&lt;p&gt;&lt;a href="https://github.com/chopratejas/headroom"&gt;https://github.com/chopratejas/headroom&lt;/a&gt;&lt;/p&gt;&#10;&lt;p&gt;&lt;strong&gt;官方的说法是它是个给 AI agent 省 token 的&amp;quot;压缩中间层&lt;/strong&gt;&amp;quot;。&lt;/p&gt;&#10;&lt;ul&gt;&#10;&lt;li&gt;AI agent(比如 Claude Code)干活时,要把一大堆东西塞给大模型读——工具输出、日志、报错、检索结果、文件内容、聊天历史。&lt;/li&gt;&#10;&lt;li&gt;这些东西又臭又长,烧token、烧钱、还容易把上下文撑爆。&lt;/li&gt;&#10;&lt;li&gt;Headroom 在这些内容到达大模型之前先压缩一遍,号称答案不变、token 砍掉 60–95%。&lt;/li&gt;&#10;&lt;/ul&gt;&#10;&lt;p&gt;怎么做到的呢？&lt;/p&gt;</description></item><item><title>agent eval：（一）deepeval</title><link>https://moyutianzun.cn/en/blog/agent-eval-yi-deepeval/</link><pubDate>Tue, 19 May 2026 03:07:56 +0800</pubDate><guid>https://moyutianzun.cn/en/blog/agent-eval-yi-deepeval/</guid><description>&lt;h1 id="特性分析"&gt;特性分析&lt;a class="hl-a" href="#%e7%89%b9%e6%80%a7%e5%88%86%e6%9e%90" aria-hidden="true"&gt;#&lt;/a&gt;&lt;/h1&gt;&#10;&lt;h2 id="集成支持"&gt;集成支持&lt;a class="hl-a" href="#%e9%9b%86%e6%88%90%e6%94%af%e6%8c%81" aria-hidden="true"&gt;#&lt;/a&gt;&lt;/h2&gt;&#10;&lt;p&gt;deepeval 通过一个统一的 Tracing 核心层来兼容所有框架。无论外部框架的形式如何不同，最终都汇聚到同一套数据模型和 trace 管理器。&lt;/p&gt;&#10;&lt;p&gt;&lt;img src="https://moyutianzun.cn/images/20260518201906340.png" alt="" loading="lazy" decoding="async"&gt;&#10;&lt;/p&gt;&#10;&lt;ul&gt;&#10;&lt;li&gt;&lt;strong&gt;Monkey-Patch 直接替换 SDK 类方法&lt;/strong&gt;&lt;/li&gt;&#10;&lt;li&gt;&lt;strong&gt;Callback Handler 实现框架原生回调接口&lt;/strong&gt;&lt;/li&gt;&#10;&lt;li&gt;&lt;strong&gt;Trace Processor 实现框架的追踪处理器协议&lt;/strong&gt;&lt;/li&gt;&#10;&lt;li&gt;&lt;strong&gt;OTel 桥接则通过 OpenTelemetry 标准协议转换&lt;/strong&gt;&lt;/li&gt;&#10;&lt;/ul&gt;&#10;&lt;p&gt;所有路径最终都通过 TraceManager 统一管理 span 的生命周期。&lt;/p&gt;</description></item><item><title>vibe coding系列：（三）一边vibe一边read</title><link>https://moyutianzun.cn/en/blog/yi-bian-vibeyi-bian-read/</link><pubDate>Mon, 04 May 2026 09:33:24 +0800</pubDate><guid>https://moyutianzun.cn/en/blog/yi-bian-vibeyi-bian-read/</guid><description>&lt;h1 id="安装"&gt;安装&lt;a class="hl-a" href="#%e5%ae%89%e8%a3%85" aria-hidden="true"&gt;#&lt;/a&gt;&lt;/h1&gt;&#10;&lt;p&gt;powershell安装&lt;/p&gt;&#10;&lt;pre tabindex="0"&gt;&lt;code&gt;scoop bucket add extras&#10;scoop install extras/wezterm&#10;scoop install pwsh&#10;scoop install yazi&#10;&lt;/code&gt;&lt;/pre&gt;&lt;p&gt;这里的pwsh是powershell 7，现在更多都是用pwsh了，后面主题出现的inner liner错误，基本上都是没用pwsh的原因。&lt;/p&gt;</description></item><item><title>agent系列（五）：agent架构的落地思考</title><link>https://moyutianzun.cn/en/blog/agentxi-lie-wu-ru-he-gou-jian-hao-yi-ge-qaxi-tong/</link><pubDate>Wed, 22 Apr 2026 07:43:23 +0800</pubDate><guid>https://moyutianzun.cn/en/blog/agentxi-lie-wu-ru-he-gou-jian-hao-yi-ge-qaxi-tong/</guid><description>&lt;p&gt;最近在楼下停车场散步，反思了一下最近遇到的bug，觉得蛮有意思的，写篇blog分享一下。&lt;/p&gt;&#10;&lt;h1 id="prompt-engineering的useful和useless"&gt;prompt engineering的useful和useless&lt;a class="hl-a" href="#prompt-engineering%e7%9a%84useful%e5%92%8cuseless" aria-hidden="true"&gt;#&lt;/a&gt;&lt;/h1&gt;&#10;&lt;p&gt;上文我们已经从架构的方面去讨论过多种的agent架构，主要还是分为隐式plan、半显式plan和显式plan的三种不同的agent风格，&lt;/p&gt;</description></item><item><title>vibe coding系列：（二）GET SHIT DONE</title><link>https://moyutianzun.cn/en/blog/vibe-codingxi-lie-er-get-shit-done/</link><pubDate>Tue, 14 Apr 2026 06:31:13 +0800</pubDate><guid>https://moyutianzun.cn/en/blog/vibe-codingxi-lie-er-get-shit-done/</guid><description>&lt;blockquote&gt;&#10;&lt;p&gt;github link：https://github.com/gsd-build/get-shit-done/blob/main/README.zh-CN.md&lt;/p&gt;</description></item><item><title>agent系列：（四）session管理和agent架构</title><link>https://moyutianzun.cn/en/blog/agentxi-lie-si-agentjia-gou/</link><pubDate>Sat, 14 Mar 2026 16:24:52 +0800</pubDate><guid>https://moyutianzun.cn/en/blog/agentxi-lie-si-agentjia-gou/</guid><description>&lt;p&gt;如果要去design一个agent架构，避免不了的就是涉及到对session的管理和agent的架构，一个是表面的，一个是深层的&lt;/p&gt;&#10;&lt;h1 id="session管理"&gt;session管理&lt;a class="hl-a" href="#session%e7%ae%a1%e7%90%86" aria-hidden="true"&gt;#&lt;/a&gt;&lt;/h1&gt;&#10;&lt;p&gt;常见可以分成以下几类：&lt;/p&gt;</description></item><item><title>agent系列：（三）context and memory</title><link>https://moyutianzun.cn/en/blog/agentxi-lie-san-context/</link><pubDate>Thu, 05 Mar 2026 17:06:39 +0800</pubDate><guid>https://moyutianzun.cn/en/blog/agentxi-lie-san-context/</guid><description>&lt;p&gt;context其实跟infra息息相关，好的infra能支持非常多的idea并发的去实现，而pipeline的高度耦合使得infra变得屎山中的屎山，所以模块化的context infra势在必行。我们想构建一个context playload，一个&lt;strong&gt;有序、分层、可度量&lt;/strong&gt;的 typed blocks 集合，是模型单次推理的&lt;strong&gt;输入态&lt;/strong&gt;。&lt;/p&gt;</description></item><item><title>vibe coding系列：（一）我自用的配置（包括claude不封号稳健玩法）</title><link>https://moyutianzun.cn/en/blog/vibe-codingxi-lie-yi-wo-zi-yong-de-pei-zhi/</link><pubDate>Tue, 03 Mar 2026 04:34:13 +0800</pubDate><guid>https://moyutianzun.cn/en/blog/vibe-codingxi-lie-yi-wo-zi-yong-de-pei-zhi/</guid><description>&lt;h1 id="一workflow"&gt;一、workflow&lt;a class="hl-a" href="#%e4%b8%80workflow" aria-hidden="true"&gt;#&lt;/a&gt;&lt;/h1&gt;&#10;&lt;p&gt;我习惯的方式是opencode + gpt5.2 xhigh来吵方案，然后claude code + opus + sonnet来写代码，他们有以下几个区别：&lt;/p&gt;&#10;&lt;ul&gt;&#10;&lt;li&gt;&#10;&lt;p&gt;gpt通常有包月的套餐，一天能用60~120刀不等，中转站包月通常一个月不会超过80 rmb，是非常实用的大模型套餐&lt;/p&gt;</description></item><item><title>agent系列：（二）agent plan</title><link>https://moyutianzun.cn/en/blog/agentxi-lie-er-agent-planhe-duo-agent/</link><pubDate>Wed, 04 Feb 2026 07:12:09 +0800</pubDate><guid>https://moyutianzun.cn/en/blog/agentxi-lie-er-agent-planhe-duo-agent/</guid><description>&lt;p&gt;agent最基本的就是ReAct范式，其中最关键的就是agent plan，我们主要探讨的是单agent的plan和多agent不同范式如何plan的更好。&lt;/p&gt;</description></item><item><title>skills系列（二）：skills团队和opencode-omo</title><link>https://moyutianzun.cn/en/blog/skillsxi-lie-er-ru-he-gou-jian-ge-ren-de-kai-fa-skillstuan-dui/</link><pubDate>Thu, 22 Jan 2026 03:54:11 +0800</pubDate><guid>https://moyutianzun.cn/en/blog/skillsxi-lie-er-ru-he-gou-jian-ge-ren-de-kai-fa-skillstuan-dui/</guid><description>&lt;p&gt;由于codex至今不支持subagent，我实在等不及了，用codex开发确实没那么灵光，先用claude code组建团队，反正逻辑是一致的。这里附上一条codex不询问命令的指令：&lt;/p&gt;</description></item><item><title>agent系列（一）：实用的agent架构、应用、评估和未来</title><link>https://moyutianzun.cn/en/blog/agentzong-shu-jia-gou-ying-yong-ping-gu-he-wei-lai/</link><pubDate>Sun, 18 Jan 2026 16:41:53 +0800</pubDate><guid>https://moyutianzun.cn/en/blog/agentzong-shu-jia-gou-ying-yong-ping-gu-he-wei-lai/</guid><description>&lt;h1 id="0-数模"&gt;0. 数模&lt;a class="hl-a" href="#0-%e6%95%b0%e6%a8%a1" aria-hidden="true"&gt;#&lt;/a&gt;&lt;/h1&gt;&#10;&lt;p&gt;最基础的抽象是：环境有状态&lt;span class="math math-inline"&gt;$s_{t}$&lt;/span&gt;​，你能看到的是观测&lt;span class="math math-inline"&gt;$o_{t}$&lt;/span&gt;，你做动作&lt;span class="math math-inline"&gt;$a_{t}$&lt;/span&gt;​，环境给反馈（比如奖励/成功信号）并转移到新状态。用 (&lt;strong&gt;PO)MDP&lt;/strong&gt; 写就是：&lt;/p&gt;</description></item><item><title>paper2proj：DeepCode在模态转换之间的编排</title><link>https://moyutianzun.cn/en/blog/paper2proj-deepcodezai-mo-tai-zhuan-huan-zhi-jian-de-bian-pai/</link><pubDate>Thu, 15 Jan 2026 08:23:05 +0800</pubDate><guid>https://moyutianzun.cn/en/blog/paper2proj-deepcodezai-mo-tai-zhuan-huan-zhi-jian-de-bian-pai/</guid><description>&lt;h1 id="1-概述"&gt;1. 概述&lt;a class="hl-a" href="#1-%e6%a6%82%e8%bf%b0" aria-hidden="true"&gt;#&lt;/a&gt;&lt;/h1&gt;&#10;&lt;p&gt;DeepCode的工作是将科学论文转化为可执行代码，它提炼出来的问题包括：&lt;/p&gt;&#10;&lt;ul&gt;&#10;&lt;li&gt;&lt;strong&gt;Specification Preservation（规范保留&lt;/strong&gt;）: 论文中的信息分散且多模态，难以忠实地将这些片段化的规范映射到实现中。&lt;/li&gt;&#10;&lt;li&gt;&lt;strong&gt;Global Consistency under Partial Views（局部视角下的全局一致性&lt;/strong&gt;）: 代码库由相互依赖的模块组成，但生成通常是逐文件进行的，在有限上下文下难以维护接口、类型和不变量的全局一致性。&lt;/li&gt;&#10;&lt;li&gt;&lt;strong&gt;Completion of Underspecified Designs（不完全指定设计的补全&lt;/strong&gt;）: 论文通常只详述算法核心，而忽略实现细节和实验框架，推断这些未明确的但重要的选择极具挑战。&lt;/li&gt;&#10;&lt;li&gt;&lt;strong&gt;Executable Faithfulness（可执行性忠实度&lt;/strong&gt;）: 忠实复现需要可执行系统，而非仅是貌似合理的代码。长期生成常导致包含逻辑错误、依赖冲突和脆弱管道的代码库，阻碍端到端执行。&lt;/li&gt;&#10;&lt;/ul&gt;&#10;&lt;p&gt;其提出的解决思想也很鲜明：&lt;/p&gt;</description></item><item><title>skills系列：（一）coding未来的管中窥豹</title><link>https://moyutianzun.cn/en/blog/claude-code-skills------codingwei-lai-de-guan-zhong-kui-bao/</link><pubDate>Fri, 09 Jan 2026 10:03:10 +0800</pubDate><guid>https://moyutianzun.cn/en/blog/claude-code-skills------codingwei-lai-de-guan-zhong-kui-bao/</guid><description>&lt;h1 id="0-现有的就是最好的"&gt;0. 现有的就是最好的&lt;a class="hl-a" href="#0-%e7%8e%b0%e6%9c%89%e7%9a%84%e5%b0%b1%e6%98%af%e6%9c%80%e5%a5%bd%e7%9a%84" aria-hidden="true"&gt;#&lt;/a&gt;&lt;/h1&gt;&#10;&lt;p&gt;由于claude code的Opus和sonnet过于昂贵，其实codex也能用skills，这里推荐用codex的skills，只需要把&lt;code&gt;https://github.com/obra/superpowers/tree/main/skills&lt;/code&gt; 里所有的文件夹搬到&lt;code&gt;.codex&lt;/code&gt;的&lt;code&gt;skills&lt;/code&gt;文件夹里即可。很多人吹嘘skills是什么黑科技，在我看来其实就是两个思路交错产生的结果：&lt;/p&gt;</description></item><item><title>github源码阅读：（二）OpenCode Code agent</title><link>https://moyutianzun.cn/en/blog/githubyuan-ma-yue-du------opencode-xiang-mu-jia-gou-wen-dang/</link><pubDate>Fri, 02 Jan 2026 14:59:30 +0800</pubDate><guid>https://moyutianzun.cn/en/blog/githubyuan-ma-yue-du------opencode-xiang-mu-jia-gou-wen-dang/</guid><description>&lt;h1 id="项目介绍"&gt;&lt;strong&gt;项目介绍&lt;/strong&gt;&lt;a class="hl-a" href="#%e9%a1%b9%e7%9b%ae%e4%bb%8b%e7%bb%8d" aria-hidden="true"&gt;#&lt;/a&gt;&lt;/h1&gt;&#10;&lt;h2 id="什么是-opencode"&gt;&lt;strong&gt;什么是 OpenCode&lt;/strong&gt;&lt;a class="hl-a" href="#%e4%bb%80%e4%b9%88%e6%98%af-opencode" aria-hidden="true"&gt;#&lt;/a&gt;&lt;/h2&gt;&#10;&lt;p&gt;&lt;a href="https://github.com/sst/opencode"&gt;&lt;strong&gt;OpenCode&lt;/strong&gt;&lt;/a&gt; 是一个 100% 开源的code agent，专注于为开发者提供强大、灵活且可扩展的 AI 编程体验。与 Cursor、Copilot 等商业工具不同，OpenCode 不绑定任何特定的 LLM 提供商，支持 Claude、OpenAI、Google、本地模型等多种提供商。&lt;/p&gt;</description></item><item><title>github源码阅读：（一）LlamaIndex 上下文数据规范</title><link>https://moyutianzun.cn/en/blog/githubyuan-ma------llamaindex-shang-xia-wen-shu-ju-gui-fan/</link><pubDate>Wed, 31 Dec 2025 06:09:51 +0800</pubDate><guid>https://moyutianzun.cn/en/blog/githubyuan-ma------llamaindex-shang-xia-wen-shu-ju-gui-fan/</guid><description>&lt;h1 id="概述"&gt;概述&lt;a class="hl-a" href="#%e6%a6%82%e8%bf%b0" aria-hidden="true"&gt;#&lt;/a&gt;&lt;/h1&gt;&#10;&lt;p&gt;&lt;a href="https://github.com/run-llama/llama_index"&gt;LlamaIndex&lt;/a&gt;（原名 GPT Index）是一个开源的数据框架，专门用于构建大语言模型（LLM）应用。它解决了 LLM 的一个核心局限性：&lt;strong&gt;LLM 在训练后就无法访问私有数据&lt;/strong&gt;。LlamaIndex 通过&lt;strong&gt;检索增强生成（RAG&lt;/strong&gt;）技术，将用户的私有数据与 LLM 的生成能力无缝连接。与langchain相比，langchian是编排工具流工作流等，&lt;strong&gt;而llamaindex专注于数据格式和数据索引&lt;/strong&gt;。&lt;/p&gt;</description></item><item><title>算法 —— 基础篇：DP</title><link>https://moyutianzun.cn/en/blog/suan-fa------ji-chu-pian-dp/</link><pubDate>Mon, 29 Dec 2025 09:36:52 +0800</pubDate><guid>https://moyutianzun.cn/en/blog/suan-fa------ji-chu-pian-dp/</guid><description>&lt;p&gt;题单来自于：&lt;a href="https://ac.nowcoder.com/discuss/828697?type=101&amp;amp;order=0&amp;amp;pos=2&amp;amp;page=1&amp;amp;channel=-1&amp;amp;source_id=1"&gt;【算法进阶题单】动态规划、数据结构、图论、数学、字符串、计算几何、博弈&lt;/a&gt;&lt;/p&gt;&#10;&lt;h1 id="状态机dp"&gt;状态机DP&lt;a class="hl-a" href="#%e7%8a%b6%e6%80%81%e6%9c%badp" aria-hidden="true"&gt;#&lt;/a&gt;&lt;/h1&gt;&#10;&lt;h2 id="时间序列"&gt;时间序列&lt;a class="hl-a" href="#%e6%97%b6%e9%97%b4%e5%ba%8f%e5%88%97" aria-hidden="true"&gt;#&lt;/a&gt;&lt;/h2&gt;&#10;&lt;p&gt;&lt;img src="https://moyutianzun.cn/images/20260429163652697.png" alt="" loading="lazy" decoding="async"&gt;&#10;&lt;/p&gt;&#10;&lt;p&gt;max和mini的&lt;/p&gt;&#10;&lt;p&gt;&lt;a href="https://leetcode.cn/problems/best-time-to-buy-and-sell-stock/"&gt;&lt;strong&gt;121. 买卖股票的最佳时机&lt;/strong&gt;&lt;/a&gt;&lt;/p&gt;&#10;&lt;p&gt;&lt;a href="https://leetcode.cn/problems/best-time-to-buy-and-sell-stock-ii/"&gt;&lt;strong&gt;122. 买卖股票的最佳时机 II&lt;/strong&gt;&lt;/a&gt;（&lt;strong&gt;有神中神dp&lt;/strong&gt;）&lt;/p&gt;&#10;&lt;p&gt;PD就要看&lt;a href="https://www.bilibili.com/video/BV1ho4y1W7QK"&gt;&lt;strong&gt;买卖股票的最佳时机【基础算法精讲 21&lt;/strong&gt;】&lt;/a&gt;&lt;/p&gt;&#10;&lt;p&gt;既然有次数限制，就要在遍历的过程中记录次数&lt;/p&gt;</description></item><item><title>基本操作：（一）github is all you need</title><link>https://moyutianzun.cn/en/blog/github-dai-ma-guan-li------sourcetree/</link><pubDate>Mon, 22 Dec 2025 16:56:26 +0800</pubDate><guid>https://moyutianzun.cn/en/blog/github-dai-ma-guan-li------sourcetree/</guid><description>&lt;h1 id="码农圣地-github"&gt;码农圣地 github&lt;a class="hl-a" href="#%e7%a0%81%e5%86%9c%e5%9c%a3%e5%9c%b0-github" aria-hidden="true"&gt;#&lt;/a&gt;&lt;/h1&gt;&#10;&lt;p&gt;github众所周知，是最大的&lt;del&gt;同性交友网站&lt;/del&gt;开源项目公布网站，每个人都可以在这里share自己的项目。&lt;/p&gt;&#10;&lt;p&gt;由于open-source往往由一个团队进行开发，所有有很多开发版本，为了管理不同的开发版本之间的异同，远古linux大神linus写了git的原型来管理不同的fork。每个opensource都有一条master主线，然后有不同的分支合并到master分支上。由此也看出，git最主要的功能就是派生（fork）、合并（merge）。&lt;/p&gt;</description></item><item><title>Codex 和 CC —— terminal vibe coding tutorial</title><link>https://moyutianzun.cn/en/blog/codex-he-cc------terminal-vibe-coding-tutorial/</link><pubDate>Thu, 18 Dec 2025 07:15:59 +0800</pubDate><guid>https://moyutianzun.cn/en/blog/codex-he-cc------terminal-vibe-coding-tutorial/</guid><description>&lt;p&gt;这篇文章记录我熟悉的codex和claude code的常用方式，若有更好的方法欢迎联系我。&lt;/p&gt;&#10;&lt;h1 id="安装"&gt;安装&lt;a class="hl-a" href="#%e5%ae%89%e8%a3%85" aria-hidden="true"&gt;#&lt;/a&gt;&lt;/h1&gt;&#10;&lt;p&gt;其实无论是在哪个终端安装都可以，无非是权限的问题，在wsl或者git bash中是最好的，但是我习惯在Window的terminal安装了，直接安装即可，问题不大。&lt;/p&gt;</description></item><item><title>大佬博客 —— Orz</title><link>https://moyutianzun.cn/en/blog/da-lao-bo-ke------orz/</link><pubDate>Thu, 11 Dec 2025 07:02:09 +0800</pubDate><guid>https://moyutianzun.cn/en/blog/da-lao-bo-ke------orz/</guid><description>&lt;p&gt;旨在排列出看过或者没看过收集的高质量博客&lt;/p&gt;&#10;&lt;h1 id="苏神"&gt;苏神&lt;a class="hl-a" href="#%e8%8b%8f%e7%a5%9e" aria-hidden="true"&gt;#&lt;/a&gt;&lt;/h1&gt;&#10;&lt;p&gt;地址：&lt;a href="https://kexue.fm/"&gt;科学空间&lt;/a&gt;&lt;/p&gt;&#10;&lt;h2 id="高代"&gt;高代&lt;a class="hl-a" href="#%e9%ab%98%e4%bb%a3" aria-hidden="true"&gt;#&lt;/a&gt;&lt;/h2&gt;&#10;&lt;p&gt;&lt;a href="https://kexue.fm/archives/8453"&gt;&lt;strong&gt;从一个单位向量变换到另一个单位向量的正交矩阵&lt;/strong&gt;&lt;/a&gt;&lt;/p&gt;&#10;&lt;p&gt;&lt;a href="https://kexue.fm/archives/11072"&gt;“&lt;strong&gt;对角+低秩”三角阵的高效求逆方法&lt;/strong&gt;&lt;/a&gt;&lt;/p&gt;&#10;&lt;p&gt;&lt;a href="https://kexue.fm/archives/10847"&gt;&lt;strong&gt;矩阵的有效秩（Effective Rank&lt;/strong&gt;）&lt;/a&gt;&lt;/p&gt;&#10;&lt;p&gt;&lt;a href="https://kexue.fm/archives/10366"&gt;&lt;strong&gt;低秩近似之路（一）：伪逆&lt;/strong&gt;&lt;/a&gt;&lt;/p&gt;&#10;&lt;p&gt;&lt;a href="https://kexue.fm/archives/10407"&gt;&lt;strong&gt;低秩近似之路（二）：SVD&lt;/strong&gt;&lt;/a&gt;&lt;/p&gt;&#10;&lt;p&gt;&lt;a href="https://kexue.fm/archives/10427"&gt;&lt;strong&gt;低秩近似之路（三）：CR&lt;/strong&gt;&lt;/a&gt;&lt;/p&gt;&#10;&lt;p&gt;&lt;a href="https://kexue.fm/archives/10501"&gt;&lt;strong&gt;低秩近似之路（四）：ID&lt;/strong&gt;&lt;/a&gt;&lt;/p&gt;&#10;&lt;p&gt;&lt;a href="https://kexue.fm/archives/10662"&gt;&lt;strong&gt;低秩近似之路（五）：CUR&lt;/strong&gt;&lt;/a&gt;&lt;/p&gt;&#10;&lt;p&gt;&lt;a href="https://kexue.fm/archives/10249"&gt;&lt;strong&gt;Monarch矩阵：计算高效的稀疏型矩阵分解&lt;/strong&gt;&lt;/a&gt;&lt;/p&gt;&#10;&lt;p&gt;&lt;a href="https://kexue.fm/archives/11335"&gt;&lt;strong&gt;随机矩阵的谱范数的快速估计&lt;/strong&gt;&lt;/a&gt;&lt;/p&gt;&#10;&lt;h2 id="概率与统计"&gt;概率与统计&lt;a class="hl-a" href="#%e6%a6%82%e7%8e%87%e4%b8%8e%e7%bb%9f%e8%ae%a1" aria-hidden="true"&gt;#&lt;/a&gt;&lt;/h2&gt;&#10;&lt;p&gt;&lt;a href="https://kexue.fm/archives/9085"&gt;&lt;strong&gt;从重参数的角度看离散概率分布的构建&lt;/strong&gt;&lt;/a&gt;&lt;/p&gt;&#10;&lt;p&gt;&lt;a href="https://kexue.fm/archives/10145"&gt;&lt;strong&gt;通向概率分布之路：盘点Softmax及其替代品&lt;/strong&gt;&lt;/a&gt;&lt;/p&gt;&#10;&lt;p&gt;&lt;a href="https://kexue.fm/archives/8578"&gt;&lt;strong&gt;概率视角下的线性模型：逻辑回归有解析解吗&lt;/strong&gt;？&lt;/a&gt;&lt;/p&gt;&#10;&lt;p&gt;&lt;a href="https://kexue.fm/archives/8512"&gt;&lt;strong&gt;两个多元正态分布的KL散度、巴氏距离和W距离&lt;/strong&gt;&lt;/a&gt;&lt;/p&gt;&#10;&lt;p&gt;&lt;a href="https://kexue.fm/archives/9595"&gt;&lt;strong&gt;如何度量数据的稀疏程度&lt;/strong&gt;？&lt;/a&gt;&lt;/p&gt;&#10;&lt;p&gt;&lt;a href="https://kexue.fm/archives/9368"&gt;&lt;strong&gt;从局部到全局：语义相似度的测地线距离&lt;/strong&gt;&lt;/a&gt;&lt;/p&gt;&#10;&lt;p&gt;&lt;a href="https://kexue.fm/archives/8679"&gt;&lt;strong&gt;让人惊叹的Johnson-Lindenstrauss引理：理论篇&lt;/strong&gt;&lt;/a&gt;&lt;/p&gt;&#10;&lt;p&gt;&lt;a href="https://kexue.fm/archives/8706"&gt;&lt;strong&gt;让人惊叹的Johnson-Lindenstrauss引理：应用篇&lt;/strong&gt;&lt;/a&gt;&lt;/p&gt;&#10;&lt;p&gt;&lt;a href="https://kexue.fm/archives/9588"&gt;&lt;strong&gt;从JL引理看熵不变性Attention&lt;/strong&gt;&lt;/a&gt;&lt;/p&gt;&#10;&lt;p&gt;&lt;a href="https://kexue.fm/archives/8823"&gt;&lt;strong&gt;从熵不变性看Attention的Scale操作&lt;/strong&gt;&lt;/a&gt;&lt;/p&gt;&#10;&lt;p&gt;&lt;a href="https://kexue.fm/archives/9034"&gt;&lt;strong&gt;熵不变性Softmax的一个快速推导&lt;/strong&gt;&lt;/a&gt;&lt;/p&gt;&#10;&lt;p&gt;&lt;a href="https://kexue.fm/archives/9812"&gt;&lt;strong&gt;从梯度最大化看Attention的Scale操作&lt;/strong&gt;&lt;/a&gt;&lt;/p&gt;&#10;&lt;p&gt;&lt;a href="https://kexue.fm/archives/9768"&gt;&lt;strong&gt;随机分词浅探：从Viterbi Decoding到Viterbi Sampling&lt;/strong&gt;&lt;/a&gt;&lt;/p&gt;</description></item><item><title>context engine note</title><link>https://moyutianzun.cn/en/blog/context-engine-note/</link><pubDate>Tue, 02 Dec 2025 09:28:32 +0800</pubDate><guid>https://moyutianzun.cn/en/blog/context-engine-note/</guid><description>&lt;h1 id="context-manager"&gt;Context Manager&lt;a class="hl-a" href="#context-manager" aria-hidden="true"&gt;#&lt;/a&gt;&lt;/h1&gt;&#10;&lt;h2 id="多模态数据的处理"&gt;多模态数据的处理&lt;a class="hl-a" href="#%e5%a4%9a%e6%a8%a1%e6%80%81%e6%95%b0%e6%8d%ae%e7%9a%84%e5%a4%84%e7%90%86" aria-hidden="true"&gt;#&lt;/a&gt;&lt;/h2&gt;&#10;&lt;p&gt;现在大模型系统的对话窗口和处理数据会产生大量的上下文，后续的qa往往会用到其中一部分上下文，所以建立有效的context engine是十分必要的。&lt;/p&gt;</description></item><item><title> LLM 系统架构 note</title><link>https://moyutianzun.cn/en/blog/llm-jia-gou-note/</link><pubDate>Mon, 01 Dec 2025 16:11:21 +0800</pubDate><guid>https://moyutianzun.cn/en/blog/llm-jia-gou-note/</guid><description>&lt;p&gt;LLM目前最大的困难来自于&lt;strong&gt;预训练使用的语料是足够多的且推理能力是足够强的&lt;/strong&gt;，但是如何&lt;strong&gt;在问答中交互使其自回归输出得到一个很好的结果&lt;/strong&gt;是很困难的，这就是agent出现的原因&lt;/p&gt;</description></item><item><title>deep research design —— note</title><link>https://moyutianzun.cn/en/blog/deep-research-design------note/</link><pubDate>Mon, 01 Dec 2025 16:09:53 +0800</pubDate><guid>https://moyutianzun.cn/en/blog/deep-research-design------note/</guid><description>&lt;h1 id="参考link"&gt;参考link：&lt;a class="hl-a" href="#%e5%8f%82%e8%80%83link" aria-hidden="true"&gt;#&lt;/a&gt;&lt;/h1&gt;&#10;&lt;p&gt;&lt;a href="https://github.com/mangopy/Deep-Research-Survey/tree/main"&gt;A Systematic Survey of Deep Research —— github&lt;/a&gt;&lt;/p&gt;&#10;&lt;p&gt;&lt;a href="https://github.com/scienceaix/deepresearch"&gt;Awesome Deep Research Projects&lt;/a&gt;&lt;/p&gt;</description></item><item><title>强化学习：从模态对齐到策略优化 —— 2、从GSPO、FSPO、GEPO</title><link>https://moyutianzun.cn/en/blog/qiang-hua-xue-xi-cong-mo-tai-dui-qi-dao-ce-lue-you-hua------2-cong-gspodao-gepo/</link><pubDate>Mon, 24 Nov 2025 09:54:25 +0800</pubDate><guid>https://moyutianzun.cn/en/blog/qiang-hua-xue-xi-cong-mo-tai-dui-qi-dao-ce-lue-you-hua------2-cong-gspodao-gepo/</guid><description>&lt;h1 id="上下文归纳"&gt;上下文归纳&lt;a class="hl-a" href="#%e4%b8%8a%e4%b8%8b%e6%96%87%e5%bd%92%e7%ba%b3" aria-hidden="true"&gt;#&lt;/a&gt;&lt;/h1&gt;&#10;&lt;p&gt;正如上文所说，qwen团队发现了其存在的缺陷是优势函数&lt;span class="math math-inline"&gt;$\hat{A}_i$&lt;/span&gt;stoken-level，但是奖励函数是序列级别的，这会导致训练大型、复杂的模型（如混合专家模型MoE）和处理长序列任务时尤为突出，常常导致训练过程不稳定甚至模型崩溃。接下来几篇文章分别从不同的角度去对整体流程进行优化，首先我们先确立一下流程：&lt;/p&gt;</description></item><item><title>强化学习：从模态对齐到策略优化 —— 1、从PPO 到 GPRO</title><link>https://moyutianzun.cn/en/blog/llmce-lue-you-hua-note/</link><pubDate>Thu, 20 Nov 2025 15:35:45 +0800</pubDate><guid>https://moyutianzun.cn/en/blog/llmce-lue-you-hua-note/</guid><description>&lt;h1 id="对齐到rl"&gt;对齐到RL&lt;a class="hl-a" href="#%e5%af%b9%e9%bd%90%e5%88%b0rl" aria-hidden="true"&gt;#&lt;/a&gt;&lt;/h1&gt;&#10;&lt;p&gt;&lt;strong&gt;一个未经对齐的 LLM 可能会产生无益、有害甚至有毒的内容，或者无法理解和遵循复杂、微妙的人类指令&lt;/strong&gt;。&lt;/p&gt;&#10;&lt;p&gt;&lt;strong&gt;这一根本性问题被称为对齐（Alignment），即确保 AI 系统的目标和行为与人类的价值观和意图保持一致&lt;/strong&gt;。&lt;/p&gt;</description></item><item><title>Antigravity卡在谷歌登录环节解决方案</title><link>https://moyutianzun.cn/en/blog/antigravityqia-zai-gu-ge-deng-lu-huan-jie/</link><pubDate>Thu, 20 Nov 2025 01:33:08 +0800</pubDate><guid>https://moyutianzun.cn/en/blog/antigravityqia-zai-gu-ge-deng-lu-huan-jie/</guid><description>&lt;p&gt;首先要保证自己的IP是足够纯净的IP，开了&lt;code&gt;v2ray/clash&lt;/code&gt; 之后，点开&lt;a href="https://iplark.com/"&gt;此网站&lt;/a&gt;，就能看到IP是否是原生&lt;code&gt;IP&lt;/code&gt;且评分如何。&lt;/p&gt;&#10;&lt;p&gt;&lt;img src="https://moyutianzun.cn/images/20251120093504766.png" alt="" loading="lazy" decoding="async"&gt;&#10;&lt;/p&gt;&#10;&lt;p&gt;由于&lt;code&gt;gemini&lt;/code&gt;和&lt;code&gt;claude&lt;/code&gt;对&lt;code&gt;IP&lt;/code&gt;的要求高，建议使用原生&lt;code&gt;IP&lt;/code&gt;和高评分的节点，且长期只使用这个节点，不然可能会触发斩杀，&lt;a href="https://xn--mes358aby2apfg.com/register?code=rLYZRl11"&gt;赔钱&lt;/a&gt;有原生IP，但是知道的人太多了，用的人也太多了，上面这个IP是赔钱的。&lt;a href="https://www.v2ny.me?path=register&amp;amp;code=Ae2Z9jRW"&gt;奈云&lt;/a&gt;不知道为啥没有，但是我日常用着很舒服，网速很快。&lt;/p&gt;</description></item><item><title>生理信号大模型</title><link>https://moyutianzun.cn/en/blog/sheng-li-xin-hao-da-mo-xing/</link><pubDate>Wed, 19 Nov 2025 07:48:13 +0800</pubDate><guid>https://moyutianzun.cn/en/blog/sheng-li-xin-hao-da-mo-xing/</guid><description>&lt;p&gt;2023年至2025年的研究格局显示，通过将异构生理信号（EEG、fMRI、EMG、EOG、Gaze）与文本、视觉等高层语义模态进行深度融合，领域正在向大脑模型（Large Brain Models, LBMs）和多模态大模型（Large Multimodal Models, LMMs）迁移。&lt;/p&gt;</description></item><item><title>算法 —— 基础篇：树图、高精度、二分</title><link>https://moyutianzun.cn/en/blog/suan-fa------ji-chu-pian-di-gui-di-tui-shu-tu-gao-jing-du-er-fen-he-dp/</link><pubDate>Sat, 08 Nov 2025 17:12:49 +0800</pubDate><guid>https://moyutianzun.cn/en/blog/suan-fa------ji-chu-pian-di-gui-di-tui-shu-tu-gao-jing-du-er-fen-he-dp/</guid><description>&lt;h1 id="递归递推"&gt;递归递推&lt;a class="hl-a" href="#%e9%80%92%e5%bd%92%e9%80%92%e6%8e%a8" aria-hidden="true"&gt;#&lt;/a&gt;&lt;/h1&gt;&#10;&lt;p&gt;&lt;strong&gt;时间复杂度&lt;/strong&gt;：看for的n，取最大&lt;/p&gt;&#10;&lt;p&gt;&lt;strong&gt;空间复杂度&lt;/strong&gt;：看实际运行的时候用到了多少内存。&lt;/p&gt;&#10;&lt;blockquote&gt;&#10;&lt;p&gt;在递归算法中，每次递推都需要一个栈空间来保存调用记彔，因此在计算空间复杂度时需要计算递归栈的辅助空间。&lt;/p&gt;</description></item><item><title>迈向具有深度推理的代理RAG：LLM RAG推理系统综述</title><link>https://moyutianzun.cn/en/blog/rag-note/</link><pubDate>Thu, 06 Nov 2025 08:22:20 +0800</pubDate><guid>https://moyutianzun.cn/en/blog/rag-note/</guid><description>&lt;p&gt;rag 的三个过程：&lt;/p&gt;&#10;&lt;ul&gt;&#10;&lt;li&gt;&lt;strong&gt;检索阶段&lt;/strong&gt;：从外部知识库中提取任务相关的内容&lt;/li&gt;&#10;&lt;li&gt;&lt;strong&gt;整合阶段&lt;/strong&gt;：对检索内容进行去重、冲突解决和重新排序&lt;/li&gt;&#10;&lt;li&gt;&lt;strong&gt;生成阶段&lt;/strong&gt;：基于精选上下文进行推理以得出最终答案。&lt;/li&gt;&#10;&lt;/ul&gt;&#10;&lt;h1 id="paper-note"&gt;Paper note&lt;a class="hl-a" href="#paper-note" aria-hidden="true"&gt;#&lt;/a&gt;&lt;/h1&gt;&#10;&lt;h2 id="towards-agentic-rag-with-deep-reasoning-a-survey-of-rag-reasoning-systems-in-llms"&gt;Towards Agentic RAG with Deep Reasoning: A Survey of RAG-Reasoning Systems in LLMs&lt;a class="hl-a" href="#towards-agentic-rag-with-deep-reasoning-a-survey-of-rag-reasoning-systems-in-llms" aria-hidden="true"&gt;#&lt;/a&gt;&lt;/h2&gt;&#10;&lt;p&gt;综述围绕两个过程来讲（&lt;strong&gt;RAG ⇔ Reasoning&lt;/strong&gt;）：&lt;/p&gt;</description></item><item><title>langchain note</title><link>https://moyutianzun.cn/en/blog/agent-note/</link><pubDate>Fri, 31 Oct 2025 10:00:19 +0800</pubDate><guid>https://moyutianzun.cn/en/blog/agent-note/</guid><description>&lt;h1 id="agent范式"&gt;agent范式&lt;a class="hl-a" href="#agent%e8%8c%83%e5%bc%8f" aria-hidden="true"&gt;#&lt;/a&gt;&lt;/h1&gt;&#10;&lt;p&gt;为了更好地组织智能体的“思考”与“行动”过程，业界涌现出了多种经典的架构范式。在本章中，我们将聚焦于其中最具代表性的三种，并一步步从零实现它们：&lt;/p&gt;</description></item><item><title>未归档偶然发现的好东西</title><link>https://moyutianzun.cn/en/blog/wei-gui-dang-ou-ran-fa-xian-de-hao-dong-xi/</link><pubDate>Tue, 28 Oct 2025 01:42:56 +0800</pubDate><guid>https://moyutianzun.cn/en/blog/wei-gui-dang-ou-ran-fa-xian-de-hao-dong-xi/</guid><description>&lt;h1 id="翻译"&gt;翻译&lt;a class="hl-a" href="#%e7%bf%bb%e8%af%91" aria-hidden="true"&gt;#&lt;/a&gt;&lt;/h1&gt;&#10;&lt;p&gt;&lt;a href="https://pot-app.com/"&gt;pot翻译&lt;/a&gt; + &lt;a href="https://fanyi-api.baidu.com/"&gt;百度翻译api&lt;/a&gt;&lt;/p&gt;&#10;&lt;p&gt;&lt;a href="https://hjfy.top/"&gt;幻觉翻译&lt;/a&gt;&lt;/p&gt;&#10;&lt;h1 id="dataset"&gt;dataset&lt;a class="hl-a" href="#dataset" aria-hidden="true"&gt;#&lt;/a&gt;&lt;/h1&gt;&#10;&lt;p&gt;&lt;a href="https://tianchi.aliyun.com/dataset/210595"&gt;阿里开放式多模态AI安全评测基准 OpenMMSec&lt;/a&gt;：百万开源多模态数据集&lt;/p&gt;</description></item><item><title>Triton is all you need —— Matrix Multiplication &amp; Matrix Transpose &amp; Matrix Copy</title><link>https://moyutianzun.cn/en/blog/triton-is-all-you-need------matrix-multiplication-matrix-transpose-matrix-copy/</link><pubDate>Fri, 24 Oct 2025 09:53:04 +0800</pubDate><guid>https://moyutianzun.cn/en/blog/triton-is-all-you-need------matrix-multiplication-matrix-transpose-matrix-copy/</guid><description>&lt;p&gt;在矩阵的运算中，由于现在LLM Scaling Law，现在模型的矩阵相当的巨大。而计算单元的访存和算力有限，故此通常采用分治的思想进行并行计算，即采用分块的方式进行分块运算，这称之为tile，这里涉及到大量的程序编写的架构和编译器优化，后续我们会展开来这个说。&lt;/p&gt;</description></item><item><title>deepseek OCR —— 源码解析</title><link>https://moyutianzun.cn/en/blog/deepseek-ocr------yuan-ma-jie-xi/</link><pubDate>Tue, 21 Oct 2025 16:07:59 +0800</pubDate><guid>https://moyutianzun.cn/en/blog/deepseek-ocr------yuan-ma-jie-xi/</guid><description>&lt;p&gt;先说说我的看法，现在各种评价满天飞，我看了看README，其实很好说理解，首先肯定是OCR。&lt;/p&gt;&#10;&lt;p&gt;在README中提到两个文件，一个是对image进行OCR，一个是对PDF进行OCR。其最主要的功能就是将其文字和图片进行提取，然后改变为可执行对象。&lt;/p&gt;</description></item><item><title>CUDA profile 大全 —— nsight computer &amp; nsys &amp; pytorch</title><link>https://moyutianzun.cn/en/blog/cuda-profile/</link><pubDate>Mon, 20 Oct 2025 15:18:01 +0800</pubDate><guid>https://moyutianzun.cn/en/blog/cuda-profile/</guid><description>&lt;h1 id="cuda-api"&gt;Cuda API&lt;a class="hl-a" href="#cuda-api" aria-hidden="true"&gt;#&lt;/a&gt;&lt;/h1&gt;&#10;&lt;p&gt;创建对象：&lt;/p&gt;&#10;&lt;pre tabindex="0"&gt;&lt;code&gt;#include &amp;lt;cuda_runtime.h&amp;gt;&#10;#include &amp;lt;cuda.h&amp;gt;&#10;#include &amp;lt;iostream&amp;gt;&#10;#include &amp;lt;string&amp;gt;&#10;&#10;// 获取当前机器的GPU数量&#10;cudaError_t error_id = cudaGetDeviceCount(&amp;amp;deviceCount);&#10;&#10;for (int dev = 0; dev &amp;lt; deviceCount; ++dev) {&#10; cudaSetDevice(dev);&#10; // 初始化当前device的属性获取对象&#10; cudaDeviceProp deviceProp;&#10; cudaGetDeviceProperties(&amp;amp;deviceProp, dev);&#10;&#10; printf(&amp;#34;\nDevice %d: \&amp;#34;%s\&amp;#34;\n&amp;#34;, dev, deviceProp.name);&#10;&lt;/code&gt;&lt;/pre&gt;&lt;p&gt;拿到数据后可以查看对应feature&lt;/p&gt;</description></item><item><title>Triton is all you need —— Vector Addition &amp; Reverse Array</title><link>https://moyutianzun.cn/en/blog/triton-is-all-you-need------triton/</link><pubDate>Fri, 17 Oct 2025 08:12:03 +0800</pubDate><guid>https://moyutianzun.cn/en/blog/triton-is-all-you-need------triton/</guid><description>&lt;h1 id="vector-addition"&gt;Vector addition&lt;a class="hl-a" href="#vector-addition" aria-hidden="true"&gt;#&lt;/a&gt;&lt;/h1&gt;&#10;&lt;p&gt;&lt;strong&gt;原题&lt;/strong&gt;：编写一个在 GPU 上执行 32 位浮点数&lt;strong&gt;向量逐元素相加&lt;/strong&gt;的程序。该程序应接受两个等长输入向量，并产生一个包含它们和的输出向量。&lt;/p&gt;&#10;&lt;p&gt;exp1：&lt;/p&gt;&#10;&lt;pre tabindex="0"&gt;&lt;code class="language-auto" data-lang="auto"&gt;Input: A = [1.0, 2.0, 3.0, 4.0]&#10; B = [5.0, 6.0, 7.0, 8.0]&#10;Output: C = [6.0, 8.0, 10.0, 12.0]&#10;&lt;/code&gt;&lt;/pre&gt;&lt;p&gt;思路：&lt;/p&gt;</description></item><item><title>LLM —— attention、MHA、MQA、GQA、MLA</title><link>https://moyutianzun.cn/en/blog/llm------attention/</link><pubDate>Mon, 13 Oct 2025 03:28:44 +0800</pubDate><guid>https://moyutianzun.cn/en/blog/llm------attention/</guid><description>&lt;p&gt;文章启蒙来自苏神的&lt;a href="https://zhuanlan.zhihu.com/p/700588653"&gt;缓存与效果的极限拉扯：从MHA、MQA、GQA到MLA&lt;/a&gt;，拜谢Orz。&lt;/p&gt;&#10;&lt;h1 id="attention"&gt;attention&lt;a class="hl-a" href="#attention" aria-hidden="true"&gt;#&lt;/a&gt;&lt;/h1&gt;&#10;&lt;p&gt;首先attention的公式我们都知道如下：&lt;/p&gt;&#10;&lt;div class="math math-display"&gt;$$&#10;\mathrm{Attention}(K,Q,V)=\mathrm{softmax}(\frac{QK^{\top}}{\sqrt{d_{k}}})V&#10;$$&lt;/div&gt;&lt;p&gt;我们尝试用数学来表述清楚这个问题，首先假设输入为一条token的&lt;span class="math math-inline"&gt;$x$&lt;/span&gt;，输出为一条token的&lt;span class="math math-inline"&gt;$o$&lt;/span&gt;，结合上面的公式得到QKV矩阵：&lt;/p&gt;</description></item><item><title>算法 —— 基础篇：数组、双指针、滑动窗口、区间</title><link>https://moyutianzun.cn/en/blog/suan-fa------shu-zu-shuang-zhi-zhen-hua-dong-chuang-kou/</link><pubDate>Fri, 03 Oct 2025 08:49:53 +0800</pubDate><guid>https://moyutianzun.cn/en/blog/suan-fa------shu-zu-shuang-zhi-zhen-hua-dong-chuang-kou/</guid><description>&lt;h1 id="双向子序列也是双指针"&gt;双向子序列也是双指针&lt;a class="hl-a" href="#%e5%8f%8c%e5%90%91%e5%ad%90%e5%ba%8f%e5%88%97%e4%b9%9f%e6%98%af%e5%8f%8c%e6%8c%87%e9%92%88" aria-hidden="true"&gt;#&lt;/a&gt;&lt;/h1&gt;&#10;&lt;p&gt;&lt;a href="https://leetcode.cn/problems/trapping-rain-water/description/"&gt;接雨水&lt;/a&gt;&lt;/p&gt;&#10;&lt;p&gt;&lt;a href="https://leetcode.cn/problems/container-with-most-water/description"&gt;接水最多的容器&lt;/a&gt;&lt;/p&gt;&#10;&lt;h2 id="双指针的搜索范围"&gt;双指针的搜索范围&lt;a class="hl-a" href="#%e5%8f%8c%e6%8c%87%e9%92%88%e7%9a%84%e6%90%9c%e7%b4%a2%e8%8c%83%e5%9b%b4" aria-hidden="true"&gt;#&lt;/a&gt;&lt;/h2&gt;&#10;&lt;p&gt;&lt;a href="https://leetcode.cn/problems/3sum/description"&gt;三数之和&lt;/a&gt;&lt;/p&gt;</description></item><item><title>PD分离 —— Prefix Cache和Chunk Prefills</title><link>https://moyutianzun.cn/en/blog/pdfen-chi------prefix-cachehe-chunk-prefills/</link><pubDate>Mon, 29 Sep 2025 09:53:22 +0800</pubDate><guid>https://moyutianzun.cn/en/blog/pdfen-chi------prefix-cachehe-chunk-prefills/</guid><description>&lt;p&gt;目前推理框架基本上都需要用到多轮对话的场景，自然产生了&lt;code&gt;kv cache&lt;/code&gt;的存储和索引算法。如果能把&lt;code&gt;prompt&lt;/code&gt;和后续产生的&lt;code&gt;KV Cache&lt;/code&gt;保存下来，会极大地降低首&lt;code&gt;Token&lt;/code&gt;的耗时。&lt;/p&gt;</description></item><item><title>vllm v1 源码解析 —— 单机八卡推理</title><link>https://moyutianzun.cn/en/blog/vllm-v1-yuan-ma-jie-xi------dan-ji-ba-qia/</link><pubDate>Fri, 26 Sep 2025 09:10:35 +0800</pubDate><guid>https://moyutianzun.cn/en/blog/vllm-v1-yuan-ma-jie-xi------dan-ji-ba-qia/</guid><description>&lt;p&gt;单机八卡，我们按照PP + TP的方式来进行方案说明，使用的是vllm框架，主要命令和函数如下：&lt;/p&gt;&#10;&lt;pre tabindex="0"&gt;&lt;code&gt;python single_node_multi_gpu_demo.py --mode pipeline_parallel --tensor-parallel 4 --pipeline-parallel 2 --model facebook/opt-13b&#10;&#10;def pipeline_parallel_inference(self, model_name: str, tensor_parallel_size: int, pipeline_parallel_size: int):&#10; &amp;#34;&amp;#34;&amp;#34;流水线并行推理 - 将模型层分布到多个GPU上&amp;#34;&amp;#34;&amp;#34;&#10; print(f&amp;#34;🚀 启动流水线并行推理 - 模型: {model_name}&amp;#34;)&#10; print(f&amp;#34; 张量并行: {tensor_parallel_size}, 流水线并行: {pipeline_parallel_size}&amp;#34;)&#10; &#10; try:&#10; # 创建LLM实例，使用张量并行 + 流水线并行&#10; self.llm = LLM(&#10; model=model_name,&#10; tensor_parallel_size=tensor_parallel_size,&#10; pipeline_parallel_size=pipeline_parallel_size,&#10; # 流水线并行配置&#10; max_num_seqs=16, # 流水线并行时建议较小的batch size&#10; gpu_memory_utilization=0.8,&#10; trust_remote_code=True,&#10; )&#10; &#10; prompts = self.setup_test_prompts()&#10; print(f&amp;#34;📝 处理 {len(prompts)} 个提示词...&amp;#34;)&#10; &#10; # 批量推理&#10; start_time = time.time()&#10; outputs = self.llm.generate(prompts, self.sampling_params)&#10; end_time = time.time()&#10; &#10; # 显示结果&#10; self._display_results(outputs, end_time - start_time)&#10; &#10; except Exception as e:&#10; print(f&amp;#34;❌ 流水线并行推理失败: {e}&amp;#34;)&#10;&lt;/code&gt;&lt;/pre&gt;&lt;p&gt;在单机八卡的场景下，我们假设使用如下配置：&lt;/p&gt;</description></item><item><title>vllm v1 源码解析 —— Core</title><link>https://moyutianzun.cn/en/blog/vllm-v1-yuan-ma-jie-xi------core/</link><pubDate>Tue, 23 Sep 2025 03:56:10 +0800</pubDate><guid>https://moyutianzun.cn/en/blog/vllm-v1-yuan-ma-jie-xi------core/</guid><description>&lt;p&gt;一个client建立之后就会建立一个core engine，这些配置会通过QMZ IPC发送给core engine。&lt;/p&gt;&#10;&lt;h1 id="core-engine-architecture"&gt;Core engine Architecture&lt;a class="hl-a" href="#core-engine-architecture" aria-hidden="true"&gt;#&lt;/a&gt;&lt;/h1&gt;&#10;&lt;h2 id="worker-and-executor"&gt;Worker and Executor&lt;a class="hl-a" href="#worker-and-executor" aria-hidden="true"&gt;#&lt;/a&gt;&lt;/h2&gt;&#10;&lt;h2 id="multiprocexecutor"&gt;MultiprocExecutor&lt;a class="hl-a" href="#multiprocexecutor" aria-hidden="true"&gt;#&lt;/a&gt;&lt;/h2&gt;&#10;&lt;p&gt;在MultiprocExecutor类中，可以清晰的找到三部曲：&lt;/p&gt;</description></item><item><title>NCCL —— 标杆</title><link>https://moyutianzun.cn/en/blog/nccl------biao-gan/</link><pubDate>Mon, 22 Sep 2025 09:58:52 +0800</pubDate><guid>https://moyutianzun.cn/en/blog/nccl------biao-gan/</guid><description>&lt;p&gt;首先，需要说一些特性在前面：&lt;/p&gt;&#10;&lt;ul&gt;&#10;&lt;li&gt;NCCL是NVIDIA 集体通信库（NVIDIA Collective Communication Library），专注于 GPU 间交互，利用 NVLink、PCIe 和 InfiniBand (IB) 等互连技术实现高带宽和低延迟。&lt;/li&gt;&#10;&lt;li&gt;NCCL并不是完全开源&lt;/li&gt;&#10;&lt;li&gt;&lt;/li&gt;&#10;&lt;/ul&gt;&#10;&lt;h1 id="主要功能"&gt;主要功能&lt;a class="hl-a" href="#%e4%b8%bb%e8%a6%81%e5%8a%9f%e8%83%bd" aria-hidden="true"&gt;#&lt;/a&gt;&lt;/h1&gt;&#10;&lt;p&gt;每个GPU都可以新建一个通信器，有对应的通信Rank编号和终止操作。&lt;/p&gt;</description></item><item><title>算子进阶 —— 通信算子</title><link>https://moyutianzun.cn/en/blog/suan-zi-jin-jie------tong-xin-suan-zi/</link><pubDate>Mon, 22 Sep 2025 09:00:54 +0800</pubDate><guid>https://moyutianzun.cn/en/blog/suan-zi-jin-jie------tong-xin-suan-zi/</guid><description>&lt;p&gt;随着LLM业务的不断发展，我们发现单机单卡无法承载一个模型的训练和推理，故此出现了单机多卡和多机多卡的训练推理算子，这时候每个机和卡之间都需要通信，所以通信算子十分的重要。&lt;/p&gt;</description></item><item><title>算子进阶 —— 通算融合</title><link>https://moyutianzun.cn/en/blog/suan-zi-jin-jie------tong-suan-rong-he/</link><pubDate>Fri, 19 Sep 2025 08:16:51 +0800</pubDate><guid>https://moyutianzun.cn/en/blog/suan-zi-jin-jie------tong-suan-rong-he/</guid><description>&lt;p&gt;（施工ing）&lt;/p&gt;&#10;&lt;h1 id="概述"&gt;概述&lt;a class="hl-a" href="#%e6%a6%82%e8%bf%b0" aria-hidden="true"&gt;#&lt;/a&gt;&lt;/h1&gt;&#10;&lt;p&gt;我们知道，算子的作用是计算，那在整个体系中，我们的核心目标是拉满GPU的利用率。&lt;/p&gt;&#10;&lt;p&gt;在现代分布式体系中，多GPU之间同时存在着计算、内存访问和通信这三种基本活动，为了服务于我们的核心目标，我们需要尽可能的将通信时间和访存时间放在计算时间内，使得GPU不存在运算时间的泡泡。&lt;/p&gt;</description></item><item><title>AMD 2025 分布式推理算子优化挑战赛 —— lect 9/16 note</title><link>https://moyutianzun.cn/en/blog/amd-2025-fen-bu-shi-tui-li-suan-zi-you-hua-tiao-zhan-sai------lect-9-16-note/</link><pubDate>Fri, 19 Sep 2025 02:57:25 +0800</pubDate><guid>https://moyutianzun.cn/en/blog/amd-2025-fen-bu-shi-tui-li-suan-zi-you-hua-tiao-zhan-sai------lect-9-16-note/</guid><description>&lt;h1 id="rocm-入门"&gt;ROCm 入门&lt;a class="hl-a" href="#rocm-%e5%85%a5%e9%97%a8" aria-hidden="true"&gt;#&lt;/a&gt;&lt;/h1&gt;&#10;&lt;p&gt;&lt;img src="https://moyutianzun.cn/images/20250916190914403.png" alt="" loading="lazy" decoding="async"&gt;&#10;&lt;/p&gt;&#10;&lt;p&gt;首先就是amd官方的命名跟nv的区别，其实区别并不大，只是AMD在cuda的基础上做了更多的优化，比如说一个wavefront有64个work-item，相当于一个warp有64个threads。其次就是有两种register，在&lt;/p&gt;</description></item><item><title>Triton is all you need —— Triton 源码、编译和调试</title><link>https://moyutianzun.cn/en/blog/triton-is-all-you-need------triton-yuan-ma-bian-yi-he-diao-shi/</link><pubDate>Thu, 18 Sep 2025 07:35:21 +0800</pubDate><guid>https://moyutianzun.cn/en/blog/triton-is-all-you-need------triton-yuan-ma-bian-yi-he-diao-shi/</guid><description>&lt;p&gt;（施工ing）&lt;/p&gt;&#10;&lt;p&gt;&lt;a href="https://moyutianzun.cn/images/20250918010843097.png"&gt;&lt;img src="https://moyutianzun.cn/images/20250918010843097.png" alt="" loading="lazy" decoding="async"&gt;&#10;&lt;/a&gt;&lt;/p&gt;&#10;&lt;p&gt;&lt;img src="https://moyutianzun.cn/upload/image.png" alt="" loading="lazy" decoding="async"&gt;&#10;&lt;/p&gt;&#10;&lt;p&gt;include日录主要存放了编译器核心功能的.h头文件，提供约定和规范&lt;/p&gt;&#10;&lt;p&gt;lib是.c和.cpp，主要是功能的实现，和include一一对应&lt;/p&gt;</description></item><item><title>AI编译器 —— 笔记</title><link>https://moyutianzun.cn/en/blog/aibian-yi-qi------bi-ji/</link><pubDate>Thu, 11 Sep 2025 09:00:18 +0800</pubDate><guid>https://moyutianzun.cn/en/blog/aibian-yi-qi------bi-ji/</guid><description>&lt;h1 id="challenge"&gt;challenge&lt;a class="hl-a" href="#challenge" aria-hidden="true"&gt;#&lt;/a&gt;&lt;/h1&gt;&#10;&lt;ul&gt;&#10;&lt;li&gt;&#10;&lt;p&gt;新模型&lt;/p&gt;&#10;&lt;ul&gt;&#10;&lt;li&gt;新module出现，需要对应算子进行计算，还需要结合硬件进行特性优化和测试，尽量充分发挥硬件性能&lt;/li&gt;&#10;&lt;li&gt;硬件厂商还会发布新技术的加速计算库&lt;/li&gt;&#10;&lt;li&gt;专用加速芯片爆发导致性能可移植性成为一种刚需&lt;/li&gt;&#10;&lt;/ul&gt;&#10;&lt;/li&gt;&#10;&lt;li&gt;&#10;&lt;p&gt;不同厂商的ISA不尽相同&lt;/p&gt;</description></item><item><title>AMD 2025 分布式推理算子优化挑战赛——笔记</title><link>https://moyutianzun.cn/en/blog/amd-2025-fen-bu-shi-tui-li-suan-zi-you-hua-tiao-zhan-sai----bi-ji/</link><pubDate>Tue, 09 Sep 2025 05:48:15 +0800</pubDate><guid>https://moyutianzun.cn/en/blog/amd-2025-fen-bu-shi-tui-li-suan-zi-you-hua-tiao-zhan-sai----bi-ji/</guid><description>&lt;p&gt;比赛提供的link：&lt;/p&gt;&#10;&lt;p&gt;&lt;a href="https://modelscope.cn/competition/117/%E6%AF%94%E8%B5%9B%E7%AE%80%E4%BB%8B"&gt;魔搭社区比赛首页&lt;/a&gt; &lt;a href="https://www.datamonsters.com/amd-developer-challenge-2025"&gt;AMD比赛首页&lt;/a&gt;&lt;/p&gt;&#10;&lt;p&gt;&lt;a href="https://www.gpumode.com/v2/leaderboard/563?tab=rankings"&gt;amd-all2all kernel Leaderboard&lt;/a&gt;&lt;/p&gt;&#10;&lt;p&gt;&lt;a href="https://github.com/gpu-mode/reference-kernels/tree/main/problems/amd_distributed"&gt;reference-kernels&lt;/a&gt;&lt;/p&gt;&#10;&lt;p&gt;&lt;a href="https://discord.com/channels/"&gt;discord link&lt;/a&gt;&lt;/p&gt;&#10;&lt;p&gt;&lt;a href="https://github.com/gpu-mode/popcorn-cli?tab=readme-ov-file"&gt;Popcorn CLI&lt;/a&gt;&lt;/p&gt;&#10;&lt;p&gt;lect：&lt;/p&gt;&#10;&lt;p&gt;&lt;a href="https://www.youtube.com/watch?v=dNWv3qYU60E"&gt;ytb Bonus Lecture: AMD Developer Challenge&lt;/a&gt;&lt;/p&gt;&#10;&lt;p&gt;&lt;a href="https://stormy-sailor-96a.notion.site/Mixture-of-Experts-AMD-Problem-1d7221cc2ffa80f9b171c332aed16093"&gt;Mixture of Experts AMD Problem&lt;/a&gt;&lt;/p&gt;&#10;&lt;p&gt;&lt;a href="https://moyutianzun.cn/archives/amd-2025-fen-bu-shi-tui-li-suan-zi-you-hua-tiao-zhan-sai------lect-9-16-note"&gt;9/16 lect note&lt;/a&gt;&lt;/p&gt;</description></item><item><title>triton is all you need 之 GEMM</title><link>https://moyutianzun.cn/en/blog/triton-is-all-you-need/</link><pubDate>Sun, 07 Sep 2025 08:10:38 +0800</pubDate><guid>https://moyutianzun.cn/en/blog/triton-is-all-you-need/</guid><description>&lt;blockquote&gt;&#10;&lt;p&gt;代码参考了傅哥，请b站关注我是傅傅猪喵，谢谢喵！&lt;/p&gt;&#10;&lt;/blockquote&gt;&#10;&lt;p&gt;Triton DSL是以BLOCK tile为中心的Python DSL。与CUDA相比，Triton的使用者无法控制所有细节，因为某些优化是自动完成的，但是在Triton编译器的逐层编译优化之下也可以获得与Cuda相近甚至超过的性能。另外，Triton的编写和调试更加简单，而且学习成本更低。&lt;/p&gt;</description></item><item><title>八股系列：inference infra</title><link>https://moyutianzun.cn/en/blog/llm-infra-ba-gu-quan-ji/</link><pubDate>Wed, 03 Sep 2025 15:16:40 +0800</pubDate><guid>https://moyutianzun.cn/en/blog/llm-infra-ba-gu-quan-ji/</guid><description>&lt;p&gt;收集我和我小伙伴互相问的八股问题，里面有&lt;code&gt;gemini deep search&lt;/code&gt;的回答，望周知。&lt;/p&gt;&#10;&lt;h1 id="llm-model"&gt;LLM model&lt;a class="hl-a" href="#llm-model" aria-hidden="true"&gt;#&lt;/a&gt;&lt;/h1&gt;&#10;&lt;h2 id="rmsnorm和layernorm相比有什么优点"&gt;rmsnorm和layernorm相比有什么优点&lt;a class="hl-a" href="#rmsnorm%e5%92%8clayernorm%e7%9b%b8%e6%af%94%e6%9c%89%e4%bb%80%e4%b9%88%e4%bc%98%e7%82%b9" aria-hidden="true"&gt;#&lt;/a&gt;&lt;/h2&gt;&#10;&lt;p&gt;参考&lt;a href="https://zhuanlan.zhihu.com/p/1906310832922531627"&gt;为什么最新的大模型普遍用RMSNorm？&lt;/a&gt;&lt;/p&gt;&#10;&lt;h2 id="swiglu和relu相比有什么修改"&gt;swiglu和relu相比有什么修改&lt;a class="hl-a" href="#swiglu%e5%92%8crelu%e7%9b%b8%e6%af%94%e6%9c%89%e4%bb%80%e4%b9%88%e4%bf%ae%e6%94%b9" aria-hidden="true"&gt;#&lt;/a&gt;&lt;/h2&gt;&#10;&lt;ul&gt;&#10;&lt;li&gt;&lt;strong&gt;GELU&lt;/strong&gt;: 高斯误差线性单元&lt;/li&gt;&#10;&lt;li&gt;&lt;strong&gt;ReLU&lt;/strong&gt;: 修正线性单元&lt;/li&gt;&#10;&lt;li&gt;&lt;strong&gt;Swish (β=1.0)&lt;/strong&gt; : 标准Swish函数&lt;/li&gt;&#10;&lt;li&gt;&lt;strong&gt;Swish (β=0.5)&lt;/strong&gt; : 低β值的Swish函数&lt;/li&gt;&#10;&lt;li&gt;&lt;strong&gt;SwishLU&lt;/strong&gt;: Swish + Linear Unit组合&lt;/li&gt;&#10;&lt;/ul&gt;&#10;&lt;p&gt;其函数代码如下：&lt;/p&gt;</description></item><item><title>【CUDA从入门到入土】四、矩阵乘法</title><link>https://moyutianzun.cn/en/blog/cudacong-ru-men-dao-ru-tu-si-ju-zhen-cheng-fa/</link><pubDate>Wed, 03 Sep 2025 14:50:37 +0800</pubDate><guid>https://moyutianzun.cn/en/blog/cudacong-ru-men-dao-ru-tu-si-ju-zhen-cheng-fa/</guid><description>&lt;p&gt;矩阵乘法跟之前不同，之前一维可以直接写一个kernel，或者多个kernel线性的排布来并行计算，那么矩阵乘法就是由一维向二维转变的关键。这时候一维的kernel排布也变成了二维排布。&lt;/p&gt;</description></item><item><title>【LLM 必读综述】Speed Always Wins：LLM高效架构调查</title><link>https://moyutianzun.cn/en/blog/llm-bi-du-zong-shu-speed-always-wins-llmgao-xiao-jia-gou-diao-cha/</link><pubDate>Wed, 27 Aug 2025 08:06:02 +0800</pubDate><guid>https://moyutianzun.cn/en/blog/llm-bi-du-zong-shu-speed-always-wins-llmgao-xiao-jia-gou-diao-cha/</guid><description>&lt;p&gt;题目：Speed Always Wins: A Survey on Efficient Architectures for Large Language Models&lt;/p&gt;&#10;&lt;p&gt;作者：孙伟高 上海人工智能实验室&lt;/p&gt;&#10;&lt;p&gt;github：https://github.com/weigao266/Awesome-Efficient-Arch&lt;/p&gt;</description></item><item><title>如何最快速的找到最核心的几篇文章</title><link>https://moyutianzun.cn/en/blog/arxiv/</link><pubDate>Sun, 24 Aug 2025 16:12:48 +0800</pubDate><guid>https://moyutianzun.cn/en/blog/arxiv/</guid><description>&lt;p&gt;做这期blog的动机很简单，分享一下自己如何快速的上手某个领域的论文。最核心的三个步骤我觉得分别是：&lt;/p&gt;&#10;&lt;ul&gt;&#10;&lt;li&gt;&#10;&lt;p&gt;确定自己&lt;strong&gt;研究领域的key words&lt;/strong&gt;，这里是要从上到下的，如LLM -&amp;gt; 微调 和 CoT&lt;/p&gt;</description></item><item><title>【CUDA从入门到入土】三、reduce算子及其优化</title><link>https://moyutianzun.cn/en/blog/cudacong-ru-men-dao-ru-tu-san-reducesuan-zi-ji-qi-you-hua/</link><pubDate>Sat, 23 Aug 2025 08:46:59 +0800</pubDate><guid>https://moyutianzun.cn/en/blog/cudacong-ru-men-dao-ru-tu-san-reducesuan-zi-ji-qi-you-hua/</guid><description>&lt;p&gt;该项目代码参考&lt;a href="https://space.bilibili.com/1822828582"&gt;傅哥的课程&lt;/a&gt;，很有用的课程，请多多支持他。&lt;/p&gt;&#10;&lt;hr&gt;&#10;&lt;p&gt;reduce 规约求和是cuda中一个经典的问题，其本质是将输入的序列进行求和。&lt;/p&gt;&#10;&lt;p&gt;在CUDA的多线程中，我们清楚数据被分为一个一个的block中进行运行，每个block通过warp来并发32个线程进行运算。&lt;/p&gt;</description></item><item><title>cursor trae解决插件问题</title><link>https://moyutianzun.cn/en/blog/cursor-traejie-jue-cha-jian-wen-ti/</link><pubDate>Wed, 20 Aug 2025 12:27:24 +0800</pubDate><guid>https://moyutianzun.cn/en/blog/cursor-traejie-jue-cha-jian-wen-ti/</guid><description>&lt;h1 id="使用代理网站"&gt;使用代理网站&lt;a class="hl-a" href="#%e4%bd%bf%e7%94%a8%e4%bb%a3%e7%90%86%e7%bd%91%e7%ab%99" aria-hidden="true"&gt;#&lt;/a&gt;&lt;/h1&gt;&#10;&lt;p&gt;使用&lt;a href="https://www.vsixhub.com/"&gt;vsixhub&lt;/a&gt;进行&lt;code&gt;vsix&lt;/code&gt;下载，记得在&lt;a href="https://marketplace.visualstudio.com/vscode"&gt;marketplace&lt;/a&gt;里对比两者的版本，下载之后直接拖拽文件到&lt;code&gt;cursor&lt;/code&gt;的插件栏里就行。&lt;/p&gt;&#10;&lt;p&gt;&lt;img src="https://moyutianzun.cn/images/20250820201615828.png" alt="" loading="lazy" decoding="async"&gt;&#10;&lt;/p&gt;&#10;&lt;p&gt;&lt;img src="https://moyutianzun.cn/images/20250820201731865.png" alt="" loading="lazy" decoding="async"&gt;&#10;&lt;/p&gt;&#10;&lt;h1 id="从vscode下载"&gt;从vscode下载&lt;a class="hl-a" href="#%e4%bb%8evscode%e4%b8%8b%e8%bd%bd" aria-hidden="true"&gt;#&lt;/a&gt;&lt;/h1&gt;&#10;&lt;p&gt;从&lt;code&gt;vscode&lt;/code&gt;下载之后点击小齿轮进行下载&lt;/p&gt;&#10;&lt;p&gt;&lt;img src="https://moyutianzun.cn/images/20250820202459681.png" alt="" loading="lazy" decoding="async"&gt;&#10;&lt;/p&gt;&#10;&lt;h1 id="使用链接补充下载"&gt;使用链接补充下载&lt;a class="hl-a" href="#%e4%bd%bf%e7%94%a8%e9%93%be%e6%8e%a5%e8%a1%a5%e5%85%85%e4%b8%8b%e8%bd%bd" aria-hidden="true"&gt;#&lt;/a&gt;&lt;/h1&gt;&#10;&lt;p&gt;在下面的链接的基础上补上一部分&lt;/p&gt;&#10;&lt;pre tabindex="0"&gt;&lt;code&gt;https://marketplace.visualstudio.com/_apis/public/gallery/publishers/&#10;&lt;/code&gt;&lt;/pre&gt;&lt;p&gt;在后续补上&lt;code&gt;{发布者}/vsextensions/{插件名}/{版本号}/vspackage&lt;/code&gt;&lt;/p&gt;</description></item><item><title>【CUDA从入门到入土】二、CUDA调试和必知必会 &amp; Nsight Computer 入门</title><link>https://moyutianzun.cn/en/blog/cudacong-ru-men-dao-ru-tu-er-cudabi-zhi-bi-hui/</link><pubDate>Sun, 17 Aug 2025 17:09:02 +0800</pubDate><guid>https://moyutianzun.cn/en/blog/cudacong-ru-men-dao-ru-tu-er-cudabi-zhi-bi-hui/</guid><description>&lt;p&gt;上文中，我们运行了一个简单的cuda函数，并且一次过的将其运行了起来，这次，我们需要补充一些基础的概念，通过概念和框架的建立，我们才能走的更远，高屋建瓴的认识更多。&lt;/p&gt;</description></item><item><title>【CUDA从入门到入土】五、cuda_kernel_和_cuda_attention详解</title><link>https://moyutianzun.cn/en/blog/cudacong-ru-men-dao-ru-tu-er-cuda_kernel_he_cuda_attentionxiang-jie/</link><pubDate>Sun, 17 Aug 2025 15:14:21 +0800</pubDate><guid>https://moyutianzun.cn/en/blog/cudacong-ru-men-dao-ru-tu-er-cuda_kernel_he_cuda_attentionxiang-jie/</guid><description>&lt;p&gt;从cuda kernel出发，看懂人生第一个cuda attention&lt;/p&gt;</description></item><item><title>【CUDA从入门到入土】一、丝滑的CUDA入门</title><link>https://moyutianzun.cn/en/blog/cudacong-ru-men-dao-ru-tu-yi-si-hua-de-cudaru-men/</link><pubDate>Thu, 14 Aug 2025 16:27:54 +0800</pubDate><guid>https://moyutianzun.cn/en/blog/cudacong-ru-men-dao-ru-tu-yi-si-hua-de-cudaru-men/</guid><description>&lt;p&gt;&lt;strong&gt;CUDA是什么&lt;/strong&gt;&lt;/p&gt;&#10;&lt;p&gt;cuda是一种gpu编程组件，是一种原生支持GPU软硬件的架构，使得开发者可以直接在 GPU 上编写和执行通用计算程序。&lt;/p&gt;&#10;&lt;h2 id="gpu架构"&gt;&lt;strong&gt;GPU架构&lt;/strong&gt;&lt;a class="hl-a" href="#gpu%e6%9e%b6%e6%9e%84" aria-hidden="true"&gt;#&lt;/a&gt;&lt;/h2&gt;&#10;&lt;p&gt;&lt;img src="https://moyutianzun.cn/images/20250807194901764.png" alt="" loading="lazy" decoding="async"&gt;&#10;&lt;/p&gt;&#10;&lt;p&gt;上图是H100白皮书中，H100 GPU带满了144个SM的架构图&lt;/p&gt;</description></item><item><title>WSL2搭建cuda-triton开发环境</title><link>https://moyutianzun.cn/en/blog/wsl2%E6%90%AD%E5%BB%BAcuda-triton%E5%BC%80%E5%8F%91%E7%8E%AF%E5%A2%83/</link><pubDate>Thu, 14 Aug 2025 16:21:13 +0800</pubDate><guid>https://moyutianzun.cn/en/blog/wsl2%E6%90%AD%E5%BB%BAcuda-triton%E5%BC%80%E5%8F%91%E7%8E%AF%E5%A2%83/</guid><description>&lt;p&gt;先用了vmware + ubuntu的方法，结果发现现在gpu无法透传到vmware的虚拟机里&lt;/p&gt;&#10;&lt;p&gt;故此使用wsl2的开发环境，后续会更新许多更舒服丝滑的操作&lt;/p&gt;</description></item><item><title>【MLsys 24】keyformer阅读笔记</title><link>https://moyutianzun.cn/en/blog/wei-ming-ming-wen-zhang/</link><pubDate>Thu, 14 Aug 2025 15:54:06 +0800</pubDate><guid>https://moyutianzun.cn/en/blog/wei-ming-ming-wen-zhang/</guid><description>&lt;p&gt;文章名：Keyformer: KV Cache reduction through key tokens selection for Efficient Generative Inference&lt;/p&gt;&#10;&lt;h1 id="论文的关键"&gt;&lt;strong&gt;论文的关键&lt;/strong&gt;：&lt;a class="hl-a" href="#%e8%ae%ba%e6%96%87%e7%9a%84%e5%85%b3%e9%94%ae" aria-hidden="true"&gt;#&lt;/a&gt;&lt;/h1&gt;&#10;&lt;ol&gt;&#10;&lt;li&gt;在vllm中，Decode的时候由于矩阵乘法和KV cache的存取，该阶段受限于memory bound&lt;/li&gt;&#10;&lt;li&gt;研究发现，在推理过程中，90%的attention权重关注于特定的token&lt;/li&gt;&#10;&lt;li&gt;这篇文章的工作就是&lt;strong&gt;通过一个新颖的评分函数来找到这些特定的token&lt;/strong&gt;&lt;/li&gt;&#10;&lt;li&gt;在KV cache中就只保存这些特定的token&lt;/li&gt;&#10;&lt;li&gt;对于多embedding和多任务上，keyfromer能降低延迟，提高吞吐量，不丢失准确率&lt;/li&gt;&#10;&lt;/ol&gt;&#10;&lt;h1 id="背景"&gt;&lt;strong&gt;背景&lt;/strong&gt;&lt;a class="hl-a" href="#%e8%83%8c%e6%99%af" aria-hidden="true"&gt;#&lt;/a&gt;&lt;/h1&gt;&#10;&lt;p&gt;生成的新token会带来额外的KV cache，如MPT-7B 模型中，序列长度增加 16 倍（从 512 到 8K）会导致推理延迟增加 50 倍以上。推理总时间的约 40%（绿色突出显示）被 KV 缓存数据移动所消耗，还会延长其他操作所需的时间（蓝色显示）。当序列长度超过 8K 时，KV 缓存大小超过了模型大小。&lt;/p&gt;</description></item><item><title>MLsys24 分类汇总</title><link>https://moyutianzun.cn/en/blog/mlsys24-fen-lei-hui-zong/</link><pubDate>Thu, 14 Aug 2025 15:51:24 +0800</pubDate><guid>https://moyutianzun.cn/en/blog/mlsys24-fen-lei-hui-zong/</guid><description>&lt;h1 id="llm-推理与服务优化-llm-inference-and-serving-optimization"&gt;&lt;strong&gt;LLM 推理与服务优化 (LLM Inference and Serving Optimization)&lt;/strong&gt;&lt;a class="hl-a" href="#llm-%e6%8e%a8%e7%90%86%e4%b8%8e%e6%9c%8d%e5%8a%a1%e4%bc%98%e5%8c%96-llm-inference-and-serving-optimization" aria-hidden="true"&gt;#&lt;/a&gt;&lt;/h1&gt;&#10;&lt;h2 id="kv-缓存管理和优化-kv-cache-management-and-optimization"&gt;&lt;strong&gt;KV 缓存管理和优化 (KV Cache Management and Optimization)&lt;/strong&gt;&lt;a class="hl-a" href="#kv-%e7%bc%93%e5%ad%98%e7%ae%a1%e7%90%86%e5%92%8c%e4%bc%98%e5%8c%96-kv-cache-management-and-optimization" aria-hidden="true"&gt;#&lt;/a&gt;&lt;/h2&gt;&#10;&lt;p&gt;这些论文聚焦于 KV 缓存的减少、量化或重用，以提升生成推理效率和降低内存消耗。&lt;/p&gt;&#10;&lt;ul&gt;&#10;&lt;li&gt;&#10;&lt;p&gt;&lt;a href="http://www.moyutianzun.cn/2025/07/30/keyformer%E9%98%85%E8%AF%BB%E7%AC%94%E8%AE%B0/"&gt;Keyformer: KV Cache reduction through key tokens selection for Efficient Generative Inference&lt;/a&gt;&lt;/p&gt;</description></item><item><title>aliyun ubuntu halo AirCloud最速传说</title><link>https://moyutianzun.cn/en/blog/aliyun-ubuntu-hexo-nextzui-su-chuan-shuo/</link><pubDate>Fri, 14 Feb 2025 15:32:00 +0800</pubDate><guid>https://moyutianzun.cn/en/blog/aliyun-ubuntu-hexo-nextzui-su-chuan-shuo/</guid><description>&lt;h1 id="购买域名"&gt;购买域名&lt;a class="hl-a" href="#%e8%b4%ad%e4%b9%b0%e5%9f%9f%e5%90%8d" aria-hidden="true"&gt;#&lt;/a&gt;&lt;/h1&gt;&#10;&lt;p&gt;搜索点击阿里云域名注册服务&lt;/p&gt;&#10;&lt;p&gt;&lt;img src="https://moyutianzun.cn/images/20250729011222815.png" alt="" loading="lazy" decoding="async"&gt;&#10;&lt;/p&gt;&#10;&lt;p&gt;在此处搜索域名&lt;/p&gt;&#10;&lt;p&gt;&lt;a href="https://moyutianzun.cn/images/20250729011235979.png"&gt;&lt;img src="https://moyutianzun.cn/images/20250729011235979.png" alt="" loading="lazy" decoding="async"&gt;&#10;&lt;/a&gt;&lt;/p&gt;&#10;&lt;p&gt;英文中文都可以，然后选择购买，实名认证之后下单即可&lt;/p&gt;&#10;&lt;p&gt;下单后需要提交实名认证的模板，在域名列表-&amp;gt;解析，会被要求去实名，上传身份证，大概要几个小时，认证后即可&lt;/p&gt;</description></item></channel></rss>