<rss xmlns:atom="http://www.w3.org/2005/Atom" version="2.0"><channel><title>All Posts - Uni星人</title><link>https://unicornumr1314.github.io/posts/</link><description>All Posts | Uni星人</description><generator>Hugo -- gohugo.io</generator><language>zh-CN</language><lastBuildDate>Sat, 21 Feb 2026 10:00:00 +0800</lastBuildDate><atom:link href="https://unicornumr1314.github.io/posts/" rel="self" type="application/rss+xml"/><item><title>工具介绍：源码合并为 Word（去空行，软著专用）</title><link>https://unicornumr1314.github.io/posts/%E5%B7%A5%E5%85%B7-%E6%BA%90%E7%A0%81%E5%90%88%E5%B9%B6%E4%B8%BAword%E5%8E%BB%E7%A9%BA%E8%A1%8C/</link><pubDate>Sat, 21 Feb 2026 10:00:00 +0800</pubDate><author><name>Uni星人</name></author><guid>https://unicornumr1314.github.io/posts/%E5%B7%A5%E5%85%B7-%E6%BA%90%E7%A0%81%E5%90%88%E5%B9%B6%E4%B8%BAword%E5%8E%BB%E7%A9%BA%E8%A1%8C/</guid><description><![CDATA[<p>本工具用于将多份源代码文件合并为一个 Word 文档，并在合并过程中自动去除空行、统一格式。支持输入页眉文本（每页顶部左对齐显示），适合在申请软件著作权时快速生成“排版规范、紧凑清晰”的源代码文档。</p>
<ul>
<li>功能要点：
<ul>
<li>多文件合并（按文件名排序），可指定“起始文件”</li>
<li>去除空行（仅空白字符的行也会删除）</li>
<li>每个文件可选择是否从新页开始（分页）</li>
<li>页眉文本左对齐显示于每页顶部</li>
<li>统一格式：宋体（SimSun），字号 10 磅</li>
<li>本地处理，不上传文件</li>
</ul>
</li>
<li>适用场景：
<ul>
<li>软著申报时提交的源代码文档整理与排版</li>
<li>将分散在多个文件中的示例代码合并为统一格式的文档</li>
</ul>
</li>
<li>直接打开工具：../static/tools/源码合并为Word(去空行).html</li>
</ul>
<p>使用步骤：</p>
<ol>
<li>打开工具页面，填写“页眉文本”（如“软件著作权-源代码文档”）。</li>
<li>拖拽或选择多个源文件，工具会按文件名排序；可在“起始文件”下拉中设定从哪个文件开始。</li>
<li>根据需要勾选“文件之间插入分页”，便于分卷阅读与审阅。</li>
<li>点击“生成 Word 文档”，完成后点击“下载生成的文档”保存。</li>
</ol>
<p>提示：工具在浏览器端完成处理，适合隐私与合规要求较高的场景。*** End Patch***}ETwitterозяется. There was an error. Please try again.***}``` &ndash;&gt;</p>]]></description></item><item><title>工具介绍：ipynb 转 md（网页版）</title><link>https://unicornumr1314.github.io/posts/%E5%B7%A5%E5%85%B7-ipynb%E8%BD%AC%E5%8C%96%E4%B8%BAmd%E7%BD%91%E9%A1%B5%E7%89%88/</link><pubDate>Sun, 30 Nov 2025 10:10:00 +0800</pubDate><author><name>Uni星人</name></author><guid>https://unicornumr1314.github.io/posts/%E5%B7%A5%E5%85%B7-ipynb%E8%BD%AC%E5%8C%96%E4%B8%BAmd%E7%BD%91%E9%A1%B5%E7%89%88/</guid><description><![CDATA[<p>本工具用于将 Jupyter Notebook（<code>.ipynb</code>）内容转化为 Markdown（<code>.md</code>）格式的网页版本，便于分享与发布。</p>
<ul>
<li>功能要点：浏览器端解析、生成 Markdown 文本</li>
<li>适用场景：将 Notebook 内容迁移到博客/文档系统</li>
<li><a href="/tools/ipynb%e8%bd%ac%e5%8c%96%e4%b8%bamd%28%e7%bd%91%e9%a1%b5%e7%89%88%29.html" rel="">直接打开工具</a></li>
</ul>
<p>使用方式：</p>
<ol>
<li>打开工具页面，导入/粘贴 Notebook 内容</li>
<li>一键生成 Markdown 文本，复制后保存为 <code>.md</code></li>
<li>可根据需要手动调整图片/代码块等细节</li>
</ol>
<p>提示：该工具纯前端实现，适合快速转换与分发。</p>]]></description></item><item><title>工具介绍：B萌动漫角色票数统计</title><link>https://unicornumr1314.github.io/posts/%E5%B7%A5%E5%85%B7-b%E8%90%8C%E5%8A%A8%E6%BC%AB%E8%A7%92%E8%89%B2%E7%A5%A8%E6%95%B0%E7%BB%9F%E8%AE%A1%E5%B7%A5%E5%85%B7/</link><pubDate>Sun, 30 Nov 2025 10:05:00 +0800</pubDate><author><name>Uni星人</name></author><guid>https://unicornumr1314.github.io/posts/%E5%B7%A5%E5%85%B7-b%E8%90%8C%E5%8A%A8%E6%BC%AB%E8%A7%92%E8%89%B2%E7%A5%A8%E6%95%B0%E7%BB%9F%E8%AE%A1%E5%B7%A5%E5%85%B7/</guid><description><![CDATA[<p>本工具按分隔符 <code>》</code> 将输入文本拆分为多组，自动提取每组中的“角色名 数字票”并进行累加统计。</p>
<ul>
<li>功能要点：分组解析、角色票数累加、错误格式提醒、结果下载</li>
<li>适用场景：B萌/角色票数统计、分组文档快速汇总</li>
<li><a href="/tools/B%e8%90%8c%e5%8a%a8%e6%bc%ab%e8%a7%92%e8%89%b2%e7%a5%a8%e6%95%b0%e7%bb%9f%e8%ae%a1%e5%b7%a5%e5%85%b7.html" rel="">直接打开工具</a></li>
</ul>
<p>使用步骤：</p>
<ol>
<li>将整段文本合并为一行，按 <code>》</code> 分割为多组</li>
<li>每组以“角色名 数字票”格式书写，如：<code>雪之下雪乃 3票</code></li>
<li>点击“确认统计（累加）”，可多次输入累加结果</li>
<li>支持将统计结果导出为 <code>txt</code></li>
</ol>
<p>示例输入：</p>
<div class="code-block highlight is-closed show-line-numbers  tw-group tw-my-2">
  <div class="
    
    tw-flex 
    tw-flex-row
    tw-flex-1 
    tw-justify-between 
    tw-w-full tw-bg-bgColor-secondary
    ">      
    <button 
      class="
        code-block-button
        tw-mx-2 
        tw-flex
        tw-flex-row
        tw-flex-1"
      aria-hidden="true">
          <div class="group-[.is-open]:tw-rotate-90 tw-transition-[transform] tw-duration-500 tw-ease-in-out print:!tw-hidden tw-w-min tw-h-min tw-my-1 tw-mx-1"><svg class="icon"
    xmlns="http://www.w3.org/2000/svg" viewBox="0 0 320 512"><!-- Font Awesome Free 5.15.4 by @fontawesome - https://fontawesome.com License - https://fontawesome.com/license/free (Icons: CC BY 4.0, Fonts: SIL OFL 1.1, Code: MIT License) --><path d="M285.476 272.971L91.132 467.314c-9.373 9.373-24.569 9.373-33.941 0l-22.667-22.667c-9.357-9.357-9.375-24.522-.04-33.901L188.505 256 34.484 101.255c-9.335-9.379-9.317-24.544.04-33.901l22.667-22.667c9.373-9.373 24.569-9.373 33.941 0L285.475 239.03c9.373 9.372 9.373 24.568.001 33.941z"/></svg></div>
          <p class="tw-select-none !tw-my-1">text</p>]]></description></item><item><title>工具介绍：世萌数据统计</title><link>https://unicornumr1314.github.io/posts/%E5%B7%A5%E5%85%B7-%E4%B8%96%E8%90%8C%E6%95%B0%E6%8D%AE%E7%BB%9F%E8%AE%A1/</link><pubDate>Sun, 30 Nov 2025 10:00:00 +0800</pubDate><author><name>Uni星人</name></author><guid>https://unicornumr1314.github.io/posts/%E5%B7%A5%E5%85%B7-%E4%B8%96%E8%90%8C%E6%95%B0%E6%8D%AE%E7%BB%9F%E8%AE%A1/</guid><description><![CDATA[<p>本工具用于统计文本中被 <code>[[...]]</code> 标记的内容出现频率，并支持检测 <code>Voter Id</code> 重复。</p>
<ul>
<li>功能要点：统计 <code>[[标签]]</code> 次数，检测并提醒重复的 <code>Voter Id</code></li>
<li>适用场景：世萌/票选类文本统计、结果汇总</li>
<li><a href="/tools/%e4%b8%96%e8%90%8c%e6%95%b0%e6%8d%ae%e7%bb%9f%e8%ae%a1.html" rel="">直接打开工具</a></li>
</ul>
<p>使用步骤：</p>
<ol>
<li>在输入框粘贴文本，包含 <code>[[标签]]</code> 格式</li>
<li>点击“确认统计”，工具会解析并汇总出现次数</li>
<li>如检测到重复 <code>Voter Id</code> 会提示并阻止累加</li>
<li>可一键下载统计结果为 <code>txt</code> 文件</li>
</ol>
<p>示例输入：</p>
<div class="code-block highlight is-closed show-line-numbers  tw-group tw-my-2">
  <div class="
    
    tw-flex 
    tw-flex-row
    tw-flex-1 
    tw-justify-between 
    tw-w-full tw-bg-bgColor-secondary
    ">      
    <button 
      class="
        code-block-button
        tw-mx-2 
        tw-flex
        tw-flex-row
        tw-flex-1"
      aria-hidden="true">
          <div class="group-[.is-open]:tw-rotate-90 tw-transition-[transform] tw-duration-500 tw-ease-in-out print:!tw-hidden tw-w-min tw-h-min tw-my-1 tw-mx-1"><svg class="icon"
    xmlns="http://www.w3.org/2000/svg" viewBox="0 0 320 512"><!-- Font Awesome Free 5.15.4 by @fontawesome - https://fontawesome.com License - https://fontawesome.com/license/free (Icons: CC BY 4.0, Fonts: SIL OFL 1.1, Code: MIT License) --><path d="M285.476 272.971L91.132 467.314c-9.373 9.373-24.569 9.373-33.941 0l-22.667-22.667c-9.357-9.357-9.375-24.522-.04-33.901L188.505 256 34.484 101.255c-9.335-9.379-9.317-24.544.04-33.901l22.667-22.667c9.373-9.373 24.569-9.373 33.941 0L285.475 239.03c9.373 9.372 9.373 24.568.001 33.941z"/></svg></div>
          <p class="tw-select-none !tw-my-1">text</p>]]></description></item><item><title>MDMetaGenerator：Markdown 元信息创建生成器（带 UI）</title><link>https://unicornumr1314.github.io/posts/%E4%BD%9C%E5%93%81%E4%BB%8B%E7%BB%8D-mdmetagenerator/</link><pubDate>Sun, 23 Nov 2025 00:00:00 +0000</pubDate><author><name>Uni星人</name></author><guid>https://unicornumr1314.github.io/posts/%E4%BD%9C%E5%93%81%E4%BB%8B%E7%BB%8D-mdmetagenerator/</guid><description><![CDATA[<div class="featured-image">
                <img src="/posts/image/mdmetagenerator-preview.png" referrerpolicy="no-referrer">
            </div><h2 id="摘要" class="headerLink">
    <a href="#%e6%91%98%e8%a6%81" class="header-mark"></a>1 摘要</h2><p>MDMetaGenerator 是一款面向内容创作者的桌面端工具，提供图形化界面，表单化生成 Markdown 文章的 Front Matter、摘要与正文分隔（`</p>]]></description></item><item><title>删除-docx-空行工具脚本</title><link>https://unicornumr1314.github.io/posts/%E5%88%A0%E9%99%A4-docx-%E7%A9%BA%E8%A1%8C/</link><pubDate>Wed, 12 Nov 2025 21:16:18 +0800</pubDate><author><name>Uni星人</name></author><guid>https://unicornumr1314.github.io/posts/%E5%88%A0%E9%99%A4-docx-%E7%A9%BA%E8%A1%8C/</guid><description><![CDATA[<p>本文介绍仓库内的脚本<a href="/tools/dele_null_line.py" rel="">源代码文件</a>，它可以自动删除 <code>.docx</code> 文档中的空段落，并在非空段落中合并连续换行（将多个连续换行替换为一个）。脚本同时会处理表格单元格内的段落。</p>
<p>本工具最初用于在申请软件著作权时，批量清理源代码文档中的空行，让提交给软著的源代码文档更加紧凑、排版整齐，方便审核和归档。</p>
<p>如果你不方便在本地安装 Python，也可以使用浏览器版工具：<a href="../static/tools/dele_null_line.html" rel="">删除 DOCX 空行（网页版）</a>，直接在网页中选择 <code>.docx</code> 文件并下载清理后的文档。</p>
<p>主要功能</p>
<ul>
<li>删除文档正文中空的段落（完全没有可见字符的段落）。</li>
<li>在非空段落中把连续的 <code>\n</code> 合并为单个 <code>\n</code>，避免多余的空行或换行符导致格式混乱。</li>
<li>遍历文档内的表格，按单元格同样规则清理空段落并合并换行。</li>
<li>提供简单的 Tkinter GUI，交互选择输入/输出文件路径，适合非命令行用户。</li>
</ul>
<p>依赖</p>
<ul>
<li>使用 <code>python-docx</code>（包名 <code>python-docx</code>）来读写 Word 文档。仓库中的脚本示例使用 <code>from docx import Document</code>。</li>
</ul>
<p>安装</p>
<pre><code>pip install python-docx
</code></pre>
<p>（如果要使用 live GUI，请确保当前 Python 环境包含 Tkinter。Windows 上自带标准 Python 通常包含 Tkinter。）</p>
<p>如何使用</p>
<ul>
<li>
<p>图形界面（推荐，脚本默认行为）:</p>
<ul>
<li>
<p>运行脚本：</p>
<p>python 原始内容文档/dele_null_line.py</p>
</li>
<li>
<p>脚本会弹出文件选择对话框，请选择要处理的 <code>.docx</code> 文件，并指定输出文件名。</p>
</li>
</ul>
</li>
<li>
<p>作为模块调用（在其他脚本中复用函数）:</p>
<p>from 原始内容文档.dele_null_line import remove_empty_lines_from_docx
remove_empty_lines_from_docx(&lsquo;input.docx&rsquo;, &lsquo;output.docx&rsquo;)</p>
</li>
</ul>
<p>实现要点（供维护者参考）</p>
<ul>
<li>段落处理：脚本先遍历 <code>doc.paragraphs</code>，把纯空的段落标记删除，非空段落使用正则 <code>re.sub(r&quot;\n+&quot;,&quot;\n&quot;, text)</code> 合并连续换行。</li>
<li>表格处理：进入每个 <code>doc.tables</code>，对每个单元格的 <code>cell.paragraphs</code> 做同样的清理与合并。</li>
<li>删除操作：对要删除的段落或单元格段落，使用底层 XML 节点 <code>._element.getparent().remove(...)</code> 执行删除，避免改变段落集合时引发索引问题（通过反向索引删除）。</li>
</ul>
<p>示例场景</p>]]></description></item><item><title>XVPN: Remote LAN for Creative Work</title><link>https://unicornumr1314.github.io/posts/xvpn/</link><pubDate>Sun, 26 Oct 2025 21:07:19 +0800</pubDate><author><name>Uni星人</name></author><guid>https://unicornumr1314.github.io/posts/xvpn/</guid><description><![CDATA[<p>How I use an XVPN gateway to map a distant NAS as a local drive and collaborate as if on-site.</p>
<ul>
<li>IPv6 direct connection, no relay, near–LAN speed</li>
<li>Map NAS as network drive; file transfer saturates home bandwidth</li>
<li>Low-latency review inside the same “virtual LAN”</li>
<li>L2 features help reach machines as if on the same switch</li>
</ul>
<p>This article mirrors the Chinese version and summarizes the key takeaways for English readers.</p>]]></description></item><item><title>番剧《间谍过家家》的评论分析</title><link>https://unicornumr1314.github.io/posts/%E7%95%AA%E5%89%A7%E9%97%B4%E8%B0%8D%E8%BF%87%E5%AE%B6%E5%AE%B6%E7%9A%84%E8%AF%84%E8%AE%BA%E5%88%86%E6%9E%90/</link><pubDate>Mon, 11 Aug 2025 21:19:41 +0800</pubDate><author><name>Uni星人</name></author><guid>https://unicornumr1314.github.io/posts/%E7%95%AA%E5%89%A7%E9%97%B4%E8%B0%8D%E8%BF%87%E5%AE%B6%E5%AE%B6%E7%9A%84%E8%AF%84%E8%AE%BA%E5%88%86%E6%9E%90/</guid><description><![CDATA[<h3 id="分析目的和项目背景" class="headerLink">
    <a href="#%e5%88%86%e6%9e%90%e7%9b%ae%e7%9a%84%e5%92%8c%e9%a1%b9%e7%9b%ae%e8%83%8c%e6%99%af" class="header-mark"></a>0.1 分析目的和项目背景</h3><p>目的:分析间谍过家家的评价,了解用户对该番剧的评价
,
,数据源:kaggle上的Bilibili Spy X Family数据集</p>
<h3 id="数据准备" class="headerLink">
    <a href="#%e6%95%b0%e6%8d%ae%e5%87%86%e5%a4%87" class="header-mark"></a>0.2 数据准备</h3><div class="code-block highlight is-closed show-line-numbers  tw-group tw-my-2">
  <div class="
    
    tw-flex 
    tw-flex-row
    tw-flex-1 
    tw-justify-between 
    tw-w-full tw-bg-bgColor-secondary
    ">      
    <button 
      class="
        code-block-button
        tw-mx-2 
        tw-flex
        tw-flex-row
        tw-flex-1"
      aria-hidden="true">
          <div class="group-[.is-open]:tw-rotate-90 tw-transition-[transform] tw-duration-500 tw-ease-in-out print:!tw-hidden tw-w-min tw-h-min tw-my-1 tw-mx-1"><svg class="icon"
    xmlns="http://www.w3.org/2000/svg" viewBox="0 0 320 512"><!-- Font Awesome Free 5.15.4 by @fontawesome - https://fontawesome.com License - https://fontawesome.com/license/free (Icons: CC BY 4.0, Fonts: SIL OFL 1.1, Code: MIT License) --><path d="M285.476 272.971L91.132 467.314c-9.373 9.373-24.569 9.373-33.941 0l-22.667-22.667c-9.357-9.357-9.375-24.522-.04-33.901L188.505 256 34.484 101.255c-9.335-9.379-9.317-24.544.04-33.901l22.667-22.667c9.373-9.373 24.569-9.373 33.941 0L285.475 239.03c9.373 9.372 9.373 24.568.001 33.941z"/></svg></div>
          <p class="tw-select-none !tw-my-1">python</p>]]></description></item><item><title>在Excel中数据加工的好用方法</title><link>https://unicornumr1314.github.io/posts/%E5%9C%A8excel%E4%B8%AD%E6%95%B0%E6%8D%AE%E5%8A%A0%E5%B7%A5%E7%9A%84%E5%A5%BD%E7%94%A8%E6%96%B9%E6%B3%95/</link><pubDate>Sat, 14 Jun 2025 21:07:19 +0800</pubDate><author><name>Uni星人</name></author><guid>https://unicornumr1314.github.io/posts/%E5%9C%A8excel%E4%B8%AD%E6%95%B0%E6%8D%AE%E5%8A%A0%E5%B7%A5%E7%9A%84%E5%A5%BD%E7%94%A8%E6%96%B9%E6%B3%95/</guid><description><![CDATA[<h3 id="1-项目概览" class="headerLink">
    <a href="#1-%e9%a1%b9%e7%9b%ae%e6%a6%82%e8%a7%88" class="header-mark"></a>0.5 <strong>1. 项目概览</strong></h3><ul>
<li><strong>项目背景</strong>：某公司开展员工满意度调查，回收问卷后得到原始数据，需要对原始数据进行清洗.</li>
<li><strong>数据规模</strong>：标注数据量1000条数据+6个维度、数据来源企业脱敏数据。</li>
<li><strong>项目周期</strong>：独立完成，耗时4h。</li>
</ul>
<h4 id="2-技术路径" class="headerLink">
    <a href="#2-%e6%8a%80%e6%9c%af%e8%b7%af%e5%be%84" class="header-mark"></a>0.5.1 <strong>2. 技术路径</strong></h4><ul>
<li><strong>数据处理</strong>：经过分析,该原始数据存在以下问题：
<ol>
<li>部分问卷的 “年龄” 字段为空值；</li>
<li>“入职日期” 字段格式不一致（如 “2023-01”“2023 年 1 月”“2023/1”）；</li>
<li>存在重复提交的问卷（同一员工提交了 2 次）；</li>
</ol>
</li>
<li><strong>处理方法</strong>：
<ol>
<li>插入数据透视表,以&quot;部门&quot;为行,以&quot;年龄为值&quot;统计每个部门平均年龄;
用Excel的筛选功能,筛选出&quot;年龄&quot;为空的行,再筛选出“部门”为“销售部”的行,统一填入销售部的平均年龄。如此可以快速填充缺失值。</li>
<li>通过 “拆分格式→文本转换→汇总统一→转日期类型” 的流程，将混乱的多格式日期，逐步规范为统一的日期值：
拆分格式：为每种日期格式建列，明确处理对象；
文本转换：用 IF+MID 把不同格式转成统一文本（如 YYYY-MM-DD 文本）；
汇总统一：用 IF 合并结果，确保每行一个统一文本；
转日期：用 DATE+MID 将文本转为真正的日期类型（可参与日期计算）。
（注：步骤 2 中 IF 返回的 0 需注意过滤，避免影响后续转换；函数操作在 Excel、Google Sheets 等工具中通用，语法略有差异时需调整。）</li>
<li>用Excel的“删除重复项”功能，将重复提交的问卷删除。</li>
</ol>
</li>
</ul>
<h4 id="以下是统一多格式日期数据的逐点解析" class="headerLink">
    <a href="#%e4%bb%a5%e4%b8%8b%e6%98%af%e7%bb%9f%e4%b8%80%e5%a4%9a%e6%a0%bc%e5%bc%8f%e6%97%a5%e6%9c%9f%e6%95%b0%e6%8d%ae%e7%9a%84%e9%80%90%e7%82%b9%e8%a7%a3%e6%9e%90" class="header-mark"></a>0.5.2 以下是统一多格式日期数据的逐点解析：</h4><p>1️⃣. 对每一种日期格式添加一个列
目的：先梳理数据中存在的日期格式（如 YYYY-MM-DD、MM/DD/YYYY、DD.MM.YYYY 等），为每种格式单独建列。
作用：方便后续针对不同格式的字符串，分别提取年、月、日信息。</p>
<p>2️⃣. 使用 IF 结合 MID，转化为目标格式的常规文本（false 忽略会返回 0）
逻辑：</p>
<div class="code-block highlight is-closed show-line-numbers  tw-group tw-my-2">
  <div class="
    
    tw-flex 
    tw-flex-row
    tw-flex-1 
    tw-justify-between 
    tw-w-full tw-bg-bgColor-secondary
    ">      
    <button 
      class="
        code-block-button
        tw-mx-2 
        tw-flex
        tw-flex-row
        tw-flex-1"
      aria-hidden="true">
          <div class="group-[.is-open]:tw-rotate-90 tw-transition-[transform] tw-duration-500 tw-ease-in-out print:!tw-hidden tw-w-min tw-h-min tw-my-1 tw-mx-1"><svg class="icon"
    xmlns="http://www.w3.org/2000/svg" viewBox="0 0 320 512"><!-- Font Awesome Free 5.15.4 by @fontawesome - https://fontawesome.com License - https://fontawesome.com/license/free (Icons: CC BY 4.0, Fonts: SIL OFL 1.1, Code: MIT License) --><path d="M285.476 272.971L91.132 467.314c-9.373 9.373-24.569 9.373-33.941 0l-22.667-22.667c-9.357-9.357-9.375-24.522-.04-33.901L188.505 256 34.484 101.255c-9.335-9.379-9.317-24.544.04-33.901l22.667-22.667c9.373-9.373 24.569-9.373 33.941 0L285.475 239.03c9.373 9.372 9.373 24.568.001 33.941z"/></svg></div>
          <p class="tw-select-none !tw-my-1">text</p>]]></description></item><item><title>在Excel中清洗数据的好用方法2</title><link>https://unicornumr1314.github.io/posts/%E5%9C%A8excel%E4%B8%AD%E6%B8%85%E6%B4%97%E6%95%B0%E6%8D%AE%E7%9A%84%E5%A5%BD%E7%94%A8%E6%96%B9%E6%B3%952/</link><pubDate>Sat, 14 Jun 2025 21:07:19 +0800</pubDate><author><name>Uni星人</name></author><guid>https://unicornumr1314.github.io/posts/%E5%9C%A8excel%E4%B8%AD%E6%B8%85%E6%B4%97%E6%95%B0%E6%8D%AE%E7%9A%84%E5%A5%BD%E7%94%A8%E6%96%B9%E6%B3%952/</guid><description><![CDATA[<h3 id="1-项目概览" class="headerLink">
    <a href="#1-%e9%a1%b9%e7%9b%ae%e6%a6%82%e8%a7%88" class="header-mark"></a>0.1 <strong>1. 项目概览</strong></h3><ul>
<li><strong>项目背景</strong>：某电商平台收集了用户订单数据，存在以下问题：</li>
</ul>
<ol start="2">
<li>重复数据：同一用户 ID 出现多次相同订单（订单号重复）。</li>
<li>缺失数据：部分订单的 “支付时间” 字段为空。</li>
<li>逻辑错误：“订单金额” 字段出现负数（如 - 50 元），“商品数量” 字段出现 0 件。</li>
</ol>
<ul>
<li><strong>数据规模</strong>：标注数据量:115条，7维度、数据来源:模拟数据.</li>
<li><strong>项目周期</strong>：1h</li>
</ul>
<h4 id="2-技术路径" class="headerLink">
    <a href="#2-%e6%8a%80%e6%9c%af%e8%b7%af%e5%be%84" class="header-mark"></a>0.1.1 <strong>2. 技术路径</strong></h4><ul>
<li><strong>数据处理</strong>：</li>
</ul>
<h5 id="重复数据处理" class="headerLink">
    <a href="#%e9%87%8d%e5%a4%8d%e6%95%b0%e6%8d%ae%e5%a4%84%e7%90%86" class="header-mark"></a>0.1.1.1 重复数据处理：</h5><ol>
<li>选中 “订单号” 列，使用 “数据→删除重复项”，勾选 “订单号”，删除完全重复的订单记录。</li>
</ol>
<h5 id="缺失数据处理" class="headerLink">
    <a href="#%e7%bc%ba%e5%a4%b1%e6%95%b0%e6%8d%ae%e5%a4%84%e7%90%86" class="header-mark"></a>0.1.1.2 缺失数据处理：</h5><ol>
<li>通过 “数据→筛选” 选中所有 “支付时间” 为空的单元格。</li>
<li>计数筛选出的单元格数量.方法⑴:框选一列,然后看Excel底部状态栏的统计数据;方法⑵:找一个空单元格,使用“分类汇总” 函数计数**=SUBTOTAL(3,A:A)-1**.(3 对应的名称是 COUNTA（非空单元格个数）)</li>
<li>鉴于支付时间缺失比例达13%,而又比较重要,手动填充合理时间(订单创建时间+12h(支付和创建订单的平均间隔时间))</li>
</ol>
<h5 id="逻辑错误处理" class="headerLink">
    <a href="#%e9%80%bb%e8%be%91%e9%94%99%e8%af%af%e5%a4%84%e7%90%86" class="header-mark"></a>0.1.1.3 逻辑错误处理：</h5><ol>
<li>列出可能的逻辑错误,如 “订单金额 &lt; 0” 或 “商品数量 = 0” 的单元格.</li>
<li>使用&quot;筛选&quot;功能,筛选出&quot;商品数量&lt;=0&quot;的单元格,然后删除掉这些行.</li>
<li>使用&quot;筛选&quot;功能,筛选出&quot;订单金额&lt;0&quot;的单元格.</li>
<li>核对错误数据：若为录入错误，修正为正确值（如 - 50 元改为 50 元）；若为无效数据，删除对应记录。</li>
</ol>
<h4 id="3-附件" class="headerLink">
    <a href="#3-%e9%99%84%e4%bb%b6" class="header-mark"></a>0.1.2 <strong>3. 附件</strong></h4><ul>
<li><a href="/posts/attach/%e7%94%b5%e5%95%86%e8%ae%a2%e5%8d%95%e6%a8%a1%e6%8b%9f%e6%95%b0%e6%8d%ae.xlsx" rel="">下载模拟数据</a></li>
</ul>]]></description></item></channel></rss>