教程:获取结果数据全过程(枚举 Fork → 提取 ARGO 域名 → 探测 Bad Request → 提取 UUID → 构建关系)

本教程完整复盘 D:\DevTools\Codeartswork\T 目录下一次真实数据管道的全过程,
把每一步用到的脚本、脚本核心逻辑、输入/输出文件,以及最终得到的数据结果整合在一起,
以便复现或二次开发。

📌 归档说明(2026-09-21 重建后更新):正文 §2–§10 记录的是较早一次完整运行的快照;
2026-09-21 已用 §14 内嵌脚本完成一次全量重建并验证通过,此前 §11 全部缺口均已修复,
最新实测数值见 §0.3。全部 15 个脚本已完整内嵌到 §14,复现时直接复制即可,无需重新生成。


0. 目标与总览

目标仓库:eooce/choreo-2go(一个部署在 Cloudflare 上的代理服务模板,使用 ARGO 隧道回源)。

我们要拿到什么:

  1. 枚举 eooce/choreo-2go 的全部 fork(直接 fork + 间接 fork/network members)。

  2. 从每个 fork 的 files/index.js(或根 index.js)中提取 ARGO fallback 域名:

    const ARGO_DOMAIN = process.env.ARGO_DOMAIN || 'xxx.example.com'





  3. 探测这些域名,找出返回 “Bad Request” 的域名(说明 ARGO 隧道入口可达、但回源失败,是「可用候选」)。

  4. 从同一批 fork 中提取 UUID:

    const UUID = process.env.UUID || '00000000-0000-0000-0000-000000000000'





  5. 构建 repo ↔ domain ↔ uuid 关系映射,并输出最终数据文件。

最新实测产出(2026-09-21 全量重建,脚本见 §14):

文件 内容 最新实测
forks_direct_full.json / forks.json 直接 fork 全量数据 582
forks_all.json / all_forks.txt 全量 fork(直接 + 间接 + BFS 子 fork 并集) 796
sub_forks_full.json BFS 子 fork(间接 fork 的 fork) 795
net_repo_keys.json network members 清洗后间接 fork 736
argo_results.json 93 仓库初轮 ARGO 结果 93
argo_all_results.json 全部直接 fork ARGO 结果 582
argo_indirect_results.json 间接 fork + 子 fork 补检 ARGO 结果 796
uuid_results.json 各 repo 的 UUID 351(提取 351/351)
bad_request_all.json 合并 Bad Request 域名(direct 17 / indirect 33 / overlap 17) 33
probe_indirect_results.json 间接域名分类探测 332 域名
repo_domain_uuid.json repo → domain + uuid 映射 351 条
rel_repo_uuid.json repo → uuid 映射 351 条
rel_domain_repo.json domain → repo 映射 332 条
rel_domain_uuid.json domain → uuid 映射 332 条
repo_domains_map.json repo → domain 映射 351 条
custom_domains.txt 合并自定义域名清单 332 行
vless_links.txt VLESS 链接(模板替换 UUID/域名/用户名) 351 条

旧缺口(原 §11.1)在 2026-09-21 重建中全部解决:

  • forks_all.json 空 bug 已修复(796 条,由 ⑧ fetch_sub_forks.py 收尾并集写入);
  • BOM 由各脚本统一 encoding="utf-8-sig" 规避;
  • choreo.mugongzi123.gq 已成功映射 → cys92096/choreo-2go / 986e0d08-b275-4dd3-9e75-f3094b36fa2a。

教程样本核验一致:DLLJ34/choreo0402 → choreo0402.moisturize45.com / fb577a93-1f88-48f7-bdea-1d84c6d8ba1d、
xuanyepan/choreo-2go → choreo.wto.pp.ua / c78c9984-fe96-4f66-869b-7c0bf5c4232f。


1. 目录结构

工作目录 D:\DevTools\Codeartswork\T 下与本流程相关的文件:

脚本(.py)

脚本 作用
fetch_direct_forks.py 枚举直接 fork(GitHub REST API 分页)
fetch_network_members.py 抓取 Network members 页面 HTML
analyze_net.py 分析 network 页面结构
extract_network.py 从 HTML 提取 repo key
check_indirect_forks.py 检测间接 fork 的 ARGO 域名
check_argo.py 对 93 个仓库做初轮 ARGO 检测
check_all_forks.py 检查全部 fork 的 ARGO 域名
fetch_sub_forks.py BFS 枚举子 fork(带断点续传)
check_bad_request.py 探测直接域名(Bad Request)
probe_indirect_domains.py 探测间接域名(多状态分类)
fetch_uuid.py 提取每个 repo 的 UUID
merge_final.py 收尾合并(直接+间接 400 清单)
build_relations.py 构建 repo↔domain↔uuid 关系

数据(.json / .txt / .html)

文件 说明
forks_direct_full.json / forks.json / forks_all.json 直接 fork 全量数据
all_forks.txt / direct_forks.txt / indirect_forks.txt / indirect_forks_clean.txt fork 清单
net_dump.html Network members 页面快照
network_members.txt / network_members_owners.txt 从 HTML 提取的成员
net_repo_keys.json 从 HTML 提取的 repo key
argo_results.json 93 仓库初轮 ARGO 结果
argo_all_results.json 全部 fork ARGO 结果
argo_indirect_results.json 间接 fork ARGO 结果
sub_forks_full.json / sub_forks_state.json 子 fork BFS 结果与断点状态
custom_domains.txt 自定义域名清单(333 行)
bad_request_results.json 直接域名探测结果
probe_indirect_results.json 间接域名探测结果
uuid_results.json 各 repo 的 UUID
bad_request_all.json 合并的 35 个 Bad Request 域名
repo_domain_uuid.json 等 5 个关系文件 最终关系数据

2. 阶段一:枚举直接 fork

脚本:fetch_direct_forks.py

核心机制:

  • 调用 GitHub REST API:GET /repos/eooce/choreo-2go/forks?per_page=100&page=N
  • 分页抓取,单页失败重试 5 次,带退避。
  • 输出 forks_direct_full.json,并汇总每个 fork 的 forks_count。

HTML 一边:由于匿名 API 受限,另用 fetch_network_members.py 抓取
https://github.com/eooce/choreo-2go/network/members 页面,保存为 net_dump.html,
并从中用正则提取两类信息:

  • network_members.txt:匹配 href="/owner/repo" 形式的 repo 链接
  • network_members_owners.txt:匹配 hovercard 的 data-hovercard-url 里的 owner

后续用 analyze_net.py 确认页面结构(分页容器 / repo href / hovercard),
extract_network.py 把 repo key 提取到 net_repo_keys.json。

结果对比(已知怪癖):API 枚举得 583 个 fork,GitHub 页面显示 540 个。


3. 阶段二:初轮 ARGO 域名提取(93 仓库)

脚本:check_argo.py

这是最早的探测脚本,硬编码 93 个 owner/repo,用 urllib + ThreadPoolExecutor(max_workers=12) 并发抓取。

核心正则(本流程一直沿用):

ARGO_RE = re.compile(r"ARGO_DOMAIN\s*=\s*process\.env\.ARGO_DOMAIN\s*\|\|\s*(['\"])(.*?)\1")





抓取策略(双路径 + 双分支回退):

  1. 先读 files/index.js,再读根 index.js。
  2. 每条路径先试 main 分支,404 则回退 master。

输出:argo_results.json,并在控制台打印:

  • files/index.js 中非空 fallback 的仓库数
  • 根 index.js 中非空 fallback 的仓库数
  • 空 fallback / 缺失(404/错误)计数
  • 全量明细表

4. 阶段三:全部 fork 的 ARGO 提取(含间接 fork)

  • **check_all_forks.py**:对 fork 全量清单逐个做与 check_argo.py 相同逻辑的检测,
    汇总到 argo_all_results.json。

    已知怪癖:argo_all_results.json 首条记录带 UTF-8 BOM(\ufeff1129855376/choreo2),
    读取时必须用 encoding="utf-8-sig"。

  • **check_indirect_forks.py**:对 network members 里的间接 fork 做 ARGO 检测,
    输出 argo_indirect_results.json。每个 fork 记录 files_index / root_index 两个 fallback 候选。

  • fetch_sub_forks.py:BFS 枚举子 fork(间接 fork 的 fork),带断点续传:

    • sub_forks_state.json:保存已访问/队列状态
    • sub_forks_full.json:累积结果
    • 处理匿名 API 的 403 限流(X-RateLimit-Remaining: 0 时把任务重新入队,优雅退出)。

5. 阶段四:直接域名 Bad Request 探测

脚本:check_bad_request.py

逻辑:

  1. 从 custom_domains.txt 读入全部域名。
  2. 对每个域名先 https:// 后 http:// 请求,verify=False、allow_redirects=True。
  3. 判定 bad_request:status_code == 400 或响应体包含 "Bad Request"。
  4. 结果分两档写入 bad_request_results.json:
    • bad_request:20 个真 Bad Request
    • other:其余(502 / 530 / 200 / 403 / DNS 失败 / 超时 / 不可达等)

产出:bad_request_results.json(直接清单)

  • bad_request 共 20 个:
# 域名
1 choreo-2.ncvfjhgr.gq
2 ny.free654321.dpdns.org
3 choreo.il0p9o8i.pp.ua
4 choreo.mugongzi123.gq/
5 choreo0402.moisturize45.com
6 choreo0512.njkffdigrrg.cf
7 choreo-foolproton.foollove.pp.ua
8 choreo.startupfit.tk
9 choreo.loouggee.cf
10 choreo3.mesfhd.cf
11 choreo17.rvdfhtved.tk
12 choreo71.megaslash.cf
13 chotkl.qdo0.dpdns.org
14 choreo2-0517.kgfkyfd.eu.org
15 choreo1901.beforheart.eu.org
16 choreo.asfyfj.tk
17 choreo-eu.jlovem.ggff.net
18 chore0.fo1oklow.dpdns.org
19 choreo.wto.pp.ua
20 choreo.tulyu.eu.org

注意第 4 条 choreo.mugongzi123.gq/ 带尾部斜杠,合并脚本会 rstrip("/") 规范化。


6. 阶段五:间接域名探测(多状态分类)

脚本:probe_indirect_domains.py

对 argo_indirect_results.json 中提取的 54 个唯一域名逐个探测,并按响应状态分类:

kind 含义 数量
400_BAD_REQUEST HTTP 400 或含 “Bad Request” 16
502_BACKEND_DOWN HTTP 502 部分
530_TUNNEL_DOWN HTTP 530 且含 1033 部分
530_OTHER 其他 530 部分
200_OK 正常 200 部分
DNS_FAIL getaddrinfo 失败 部分
CONN_FAIL 连接失败 部分
ERROR 其他异常 部分
UNREACHABLE HTTPS/HTTP 均不可达 部分

产出:probe_indirect_results.json,每条含 domain / scheme / status / kind / snippet / repos。

16 个间接 400_BAD_REQUEST 域名:

# 域名 来源 repo
1 ccooho.xjmxtfgrszh.tk d57kctykt7kxd/choreo-0607, xfvs325dv/choreo-0610, xrtjzetjztjsertj/cho-2025-0605
2 cho.apiauto.dpdns.org shanjiang66/ct888
3 chogfas.csahfrrety.tk rosalyg76/chork
4 choref.cctalk.dpdns.org jipe32/cheu01
5 choreo.tulyu.eu.org jighdinaa/choo-1016
6 choreo4.adcedew1.com maoyi11/choreo4
7 choreo4.mizacgx.tk fhdjlflfpp/choreo4
8 choreo5we.cooking200.dpdns.org mbvnvng/choreo5we
9 choreoa7.hotsunny5794.gq Link208s/choreoz7
10 chosertg.kamiyah.cf Deller768/ch0er0-0831
11 eucho.edmenbender.eu.org jinefe54/eucho
12 gcysua.nytvrexefegf.tk hunvtcr/fdycasua
13 hostfree.xyz888.tk jsjs2021/hostfree
14 kciya.onserviceproviderbundle.cf mkcisyt/tusjbwefyd
15 kcysbxt.hapiutility.tk Davidinb/CHOREO-0607
16 shgmr.goko-gm.com bgthyjmk/ch0er0-3022

间接结果里还混有垃圾数据,例如 sevalla8001.b.9.a.b.0.d.0.0.1.0.a.2.ip6.arpa
(Sevalla 的 IPv6 反向主机名)和 free.hr、cho.zzx.free 等,合并时被过滤。


7. 阶段六:UUID 提取

脚本:fetch_uuid.py

逻辑:对每个候选 repo,抓取 https://raw.githubusercontent.com/{repo}/main/files/index.js
(404 回退 master),用正则提取:

r"const UUID = process\.env\.UUID \|\| '([0-9a-fA-F-]{36})'"





部分脚本做了通用化,允许引号类型与空白差异。

产出:uuid_results.json(每个 repo 的 uuid 或 NOT_FOUND / error)。


8. 阶段七:收尾合并(merge_final.py)

脚本:merge_final.py(离线确定性操作,不发起任何网络请求,不撞限流)

干了 4 件事:

  1. 直接 400 清单:读 bad_request_results.json → bad_request 列表,规范化(去空白/去尾部斜杠)→ 20 个。
  2. 间接 400 清单:读 probe_indirect_results.json → 取 kind == "400_BAD_REQUEST" → 16 个。
  3. **重建 bad_request_all.json**(合法 JSON):
    • direct_count = 20
    • indirect_count = 16
    • overlap = ["choreo.tulyu.eu.org"](唯一重叠)
    • domains = 35(合并去重排序)
  4. 归并间接自定义域名到 custom_domains.txt:
    • 从 argo_indirect_results.json 收集 fallback,过滤垃圾:
      JUNK = {"", "NOT_FOUND", "NO_MATCH", "OBFUSCATED", "error"}
    • 剔除默认值 choreo.zzx.free.hr
    • 与现有 custom_domains.txt 求并集排序写回。

输出:bad_request_all.json(35 个域名,全文见下节)。


9. 阶段八:构建关系(build_relations.py)

脚本:build_relations.py

逻辑:将各 repo 的(domain, uuid)配对,剔除 BAD_DOMAINS 后生成多个视图:

  • repo_domain_uuid.json:repo → {domain, uuid}(39 条)
  • repo_domains_map.json:repo → domain(39 条)
  • rel_repo_uuid.json:repo → uuid(39 条)
  • rel_domain_repo.json:domain → repo(34 条)
  • rel_domain_uuid.json:domain → uuid(34 条)

39 条 repo→domain→uuid 全量清单(摘自 repo_domain_uuid.json):

repo domain uuid
DLLJ34/choreo0402 choreo0402.moisturize45.com fb577a93-1f88-48f7-bdea-1d84c6d8ba1d
Davidinb/CHOREO-0607 kcysbxt.hapiutility.tk 4584ee71-bc10-489d-8484-bd1a14bb5a88
Deller768/ch0er0-0831 chosertg.kamiyah.cf 84b78553-079a-471c-93ef-c2de55202559
GF4R/choreo-2go-16 choreo.loouggee.cf f53c47a9-83f0-4260-871b-6d8bb7a155db
Link208s/choreoz7 choreoa7.hotsunny5794.gq d290c4f6-1d3d-4be1-9cec-476191bff99d
ardjsrset/choreo-2go choreo-2.ncvfjhgr.gq 525476d1-9630-46f8-aada-d53c30935d71
bgthyjmk/ch0er0-3022 shgmr.goko-gm.com d6b6b9f3-453c-48fd-882f-113cec2cd2f7
cokear/choreo-2go ny.free654321.dpdns.org 8380874e-4390-43e2-be64-f3072b0a0597
cuiwenpei/choreo-2go choreo.il0p9o8i.pp.ua 20c91d22-48d3-47df-bfc7-aa2dd0e03809
d57kctykt7kxd/choreo-0607 ccooho.xjmxtfgrszh.tk 79f07039-32e0-414d-bf39-fdf95d3b524f
dockr1/choreo-0518 choreo0512.njkffdigrrg.cf 2279dc6b-4991-446f-81be-725231f14a1b
fhdjlflfpp/choreo4 choreo4.mizacgx.tk 3719508c-f3e4-4f97-a99a-2bf0b8cb20e3
foolerhuan/choreo-newgo choreo-foolproton.foollove.pp.ua 986e0d08-b275-4dd3-9e75-f3094b36fa2a
freedomman369/choreo choreo.startupfit.tk ad3e5c8d-99ff-474b-97f9-997ca26c837c
gfrdthft/choreo0612 choreo3.mesfhd.cf e356d3cf-c85b-4c48-839c-04b885e6f7e1
ghtfd8/choreo choreo17.rvdfhtved.tk 2775a1c2-82be-45da-839a-bd954bffc859
hunvtcr/fdycasua gcysua.nytvrexefegf.tk 9f171b65-2a91-4d39-ab87-6a90c1f5a170
jighdinaa/choo-1016 choreo.tulyu.eu.org 0bf272a3-3503-487a-8fa9-eadc3768e32c
jinefe54/eucho eucho.edmenbender.eu.org f0e504a0-ae80-4e45-b082-add0349c8c10
jipe32/cheu01 choref.cctalk.dpdns.org 80a991f6-4076-4c1d-be1e-0b8573aa5bd6
jsjs2021/hostfree hostfree.xyz888.tk 5d525196-8c0e-47d1-b959-9d0d22aa01dc
kdjhca/choreo-2go choreo17.rvdfhtved.tk 2775a1c2-82be-45da-839a-bd954bffc859
maoyi11/choreo4 choreo4.adcedew1.com 3dc1fc92-c0b4-4814-abff-915a4d55fb7e
mbvnvng/choreo5we choreo5we.cooking200.dpdns.org e3f3ae02-077a-4b5e-a8fe-75c6e7584daf
mkcisyt/tusjbwefyd kciya.onserviceproviderbundle.cf 2b5f4f4b-4f09-44ed-b16f-0f81baa83a03
mkiloacea/choreo-2go choreo71.megaslash.cf 07d1cbcf-1a8e-48e0-ad5c-9017565d2f1e
nbrsx/choreo choreo.loouggee.cf f53c47a9-83f0-4260-871b-6d8bb7a155db
otuuipl/choreo-by0 chotkl.qdo0.dpdns.org a2a687bf-c8c2-4438-813e-b2633250c065
rosalyg76/chork chogfas.csahfrrety.tk 9aa91e25-31f8-49f3-9f38-491ddd800b5c
rtxytcufviughi/choreo2-0517 choreo2-0517.kgfkyfd.eu.org 1ca47020-d569-46be-9ba5-d55a18bf4652
shanjiang66/ct888 cho.apiauto.dpdns.org f23a3818-d079-48d9-86eb-03ad5a531660
ssfreely3/choreo01 choreo1901.beforheart.eu.org 619f4d6f-8be9-4480-a5d7-6dbc80ccfa79
tiquetecv/21choreo choreo.asfyfj.tk d70a04bc-1f67-448c-a174-262de5c6e366
worker690/choreo2024 chore0.fo1oklow.dpdns.org 342f4b2e-ca6b-4dbc-b0d9-4f43d2799475
woshi992681784/cho2go choreo-eu.jlovem.ggff.net dd30d58e-6663-4a77-94a6-6109ea33d967
xfvs325dv/choreo-0610 ccooho.xjmxtfgrszh.tk 79f07039-32e0-414d-bf39-fdf95d3b524f
xrtjzetjztjsertj/cho-2025-0605 ccooho.xjmxtfgrszh.tk 79f07039-32e0-414d-bf39-fdf95d3b524f
xuanyepan/choreo-2go choreo.wto.pp.ua c78c9984-fe96-4f66-869b-7c0bf5c4232f
ytenvisp/fcytsbxua choreo.tulyu.eu.org 0bf272a3-3503-487a-8fa9-eadc3768e32c

同一域名被多个 fork 共享的情况(多 fork 复用同一套部署 => 同 domain 同 uuid):

  • choreo.tulyu.eu.org ← jighdinaa/choo-1016 + ytenvisp/fcytsbxua → 0bf272a3-...
  • choreo17.rvdfhtved.tk ← ghtfd8/choreo + kdjhca/choreo-2go → 2775a1c2-...
  • choreo.loouggee.cf ← GF4R/choreo-2go-16 + nbrsx/choreo → f53c47a9-...
  • ccooho.xjmxtfgrszh.tk ← 3 个 fork → 79f07039-...(最大聚合簇)

10. 最终数据:bad_request_all.json(35 个域名)

{
"total": 35,
"direct_count": 20,
"indirect_count": 16,
"overlap": ["choreo.tulyu.eu.org"],
"domains": [
"ccooho.xjmxtfgrszh.tk",
"cho.apiauto.dpdns.org",
"chogfas.csahfrrety.tk",
"chore0.fo1oklow.dpdns.org",
"choref.cctalk.dpdns.org",
"choreo-2.ncvfjhgr.gq",
"choreo-eu.jlovem.ggff.net",
"choreo-foolproton.foollove.pp.ua",
"choreo.asfyfj.tk",
"choreo.il0p9o8i.pp.ua",
"choreo.loouggee.cf",
"choreo.mugongzi123.gq",
"choreo.startupfit.tk",
"choreo.tulyu.eu.org",
"choreo.wto.pp.ua",
"choreo0402.moisturize45.com",
"choreo0512.njkffdigrrg.cf",
"choreo17.rvdfhtved.tk",
"choreo1901.beforheart.eu.org",
"choreo2-0517.kgfkyfd.eu.org",
"choreo3.mesfhd.cf",
"choreo4.adcedew1.com",
"choreo4.mizacgx.tk",
"choreo5we.cooking200.dpdns.org",
"choreo71.megaslash.cf",
"choreoa7.hotsunny5794.gq",
"chosertg.kamiyah.cf",
"chotkl.qdo0.dpdns.org",
"eucho.edmenbender.eu.org",
"gcysua.nytvrexefegf.tk",
"hostfree.xyz888.tk",
"kciya.onserviceproviderbundle.cf",
"kcysbxt.hapiutility.tk",
"ny.free654321.dpdns.org",
"shgmr.goko-gm.com"
]
}





11. 已知缺口与注意点

11.1 关系缺口:35 个 Bad Request 只映射 34 条(已解决 ✅)

  • 旧运行中 choreo.mugongzi123.gq 找不到对应仓库与 UUID(根因:forks_all.json 为空的 bug)。
  • 2026-09-21 重建已修复:forks_all.json 796 条,该域名成功映射 →
    cys92096/choreo-2go / 986e0d08-b275-4dd3-9e75-f3094b36fa2a(33 个 Bad Request 域名全部有 repo/UUID 映射)。

11.2 其他注意点(清单)

  • BOM 坑:旧版 argo_all_results.json 首条记录带 \ufeff,所有脚本统一用 encoding="utf-8-sig" 读取规避。

  • raw.githubusercontent.com(Fastly CDN)卡顿坑:本环境下 CDN 下载大文件严重超时(单文件可达数分钟)。
    解法(⑥⑦⑪ 已采用):改走 api.github.com contents API + Accept: application/vnd.github.raw 头拿原始文件内容,
    单次约 0.6s,且带 Token 后配额 5000 次/小时完全够用。

  • PowerShell 中文输出坑:Python 往 stdout 打中文再被 shell 捕获时会报编码错/破坏管道。
    解法:脚本开头 sys.stdout.reconfigure(encoding="utf-8", errors="replace");
    运行时用 python -X utf8 + $env:PYTHONIOENCODING="utf-8" + [Console]::OutputEncoding=[System.Text.Encoding]::UTF8。

  • 垃圾域名:需过滤 free.hr、*.ip6.arpa(Sevalla 主机名)、NOT_FOUND/NO_MATCH/OBFUSCATED 等哨兵值,
    并以 choreo.zzx.free.hr 为默认值剔除。各脚本内置 JUNK / valid_domain 统一处理。

  • 限流:GitHub API 对匿名请求限流(api.github.com 60 次/小时/IP)。fetch_sub_forks.py 在
    X-RateLimit-Remaining: 0 时把任务重新入队并保存状态,保证可断点续传。
    限流最直接的解法:调用之前提供 GitHub Token——认证后配额提到 5000 次/小时,基本不会撞墙:

    请求方 限流额度
    匿名(不带 Token) 60 次/小时/IP
    带 Token 5000 次/小时
    • Token 在哪申请:GitHub → Settings → Developer settings → Personal access tokens →
      Tokens (classic),公开仓库只读不需要勾任何 scope 也能读公开 fork 清单。
    • 怎么用:所有发往 api.github.com 的请求带 Authorization: token <你的token>
      (或 Bearer <token>);发往 raw.githubusercontent.com 的请求带同样的头。
    • 推荐做法:把 token 存进环境变量 GITHUB_TOKEN,脚本里 os.environ["GITHUB_TOKEN"] 读一次,
      不要写死在代码里(避免泄露)。
    • 响应头 X-RateLimit-Limit / X-RateLimit-Remaining / X-RateLimit-Reset 用于观测剩余额度,
      剩余为 0 时才需要退避/断点续传。

12. 附录:核心脚本关键片段

12.1 ARGO 域名提取(check_argo.py 核心)

ARGO_RE = re.compile(r"ARGO_DOMAIN\s*=\s*process\.env\.ARGO_DOMAIN\s*\|\|\s*(['\"])(.*?)\1")

def get_file(repo, path):
for branch in ('main', 'master'): # main → master 回退
url = f"https://raw.githubusercontent.com/{repo}/{branch}/{path}"
try:
with urllib.request.urlopen(urllib.request.Request(url,
headers={'User-Agent': 'Mozilla/5.0 (CodeArtsWork check)'}), timeout=20) as r:
return ('ok', r.read().decode('utf-8', 'replace'))
except urllib.error.HTTPError as e:
if e.code == 404:
continue
return (f'http{e.code}', str(e))
except Exception as e:
return ('err', str(e))
return ('404', '')

def check_repo(repo):
res = {'repo': repo}
st1, c1 = get_file(repo, 'files/index.js') # 首选 files/index.js
res['files_idx_fallback'] = ARGO_RE.search(c1).group(2) if st1 == 'ok' and ARGO_RE.search(c1) else None
st2, c2 = get_file(repo, 'index.js') # 次选根 index.js
res['root_idx_fallback'] = ARGO_RE.search(c2).group(2) if st2 == 'ok' and ARGO_RE.search(c2) else None
return res

with concurrent.futures.ThreadPoolExecutor(max_workers=12) as ex:
...





12.2 间接域名分类(probe_indirect_domains.py 核心)

def classify(domain):
res = {"domain": domain, "scheme": None, "status": None, "kind": None, "snippet": ""}
for scheme in ("https", "http"):
url = f"{scheme}://{domain}"
try:
r = requests.get(url, headers=HEADERS, timeout=15, verify=False, allow_redirects=True)
status, text = r.status_code, r.text[:3000]
res.update(scheme=scheme, status=status)
if status == 400 or "Bad Request" in text:
res["kind"] = "400_BAD_REQUEST"
elif status == 502: res["kind"] = "502_BACKEND_DOWN"
elif status == 530 and "1033" in text: res["kind"] = "530_TUNNEL_DOWN"
elif status == 530: res["kind"] = "530_OTHER"
elif status == 200: res["kind"] = "200_OK"
else: res["kind"] = f"HTTP_{status}"
res["snippet"] = text[:300]
return res
except requests.exceptions.SSLError:
continue
except requests.exceptions.ConnectionError as e:
s = repr(e)
if "getaddrinfo" in s or "Name or service not known" in s or "nodename nor servname" in s:
res["kind"] = "DNS_FAIL"; res["snippet"] = "getaddrinfo failed"
else:
res["kind"] = "CONN_FAIL"; res["snippet"] = s[:200]
res["scheme"] = scheme
return res
except Exception as e:
res.update(scheme=scheme, status="ERR", kind="ERROR", snippet=repr(e)[:200])
return res
res["kind"] = "UNREACHABLE"
return res





12.3 合并规范化(merge_final.py 核心)

def normal(domain: str) -> str:
return domain.strip().rstrip("/").strip()

direct_400 = sorted({normal(i["domain"]) for i in load_json("bad_request_results.json").get("bad_request", []) if normal(i["domain"])})
indirect_400 = sorted({normal(i["domain"]) for i in load_json("probe_indirect_results.json") if i.get("kind") == "400_BAD_REQUEST"})
overlap = sorted(set(direct_400) & set(indirect_400))
merged = sorted(set(direct_400) | set(indirect_400))





12.4 带 GitHub Token 的请求模板(解决限流)

import os
import urllib.request

TOKEN = os.environ.get("GITHUB_TOKEN", "") # 调用前先设置,避免写死
HEADERS = {
"User-Agent": "Mozilla/5.0 (CodeArtsWork collect)",
"Accept": "application/vnd.github+json",
}
if TOKEN:
HEADERS["Authorization"] = f"Bearer {TOKEN}" # 也可用 f"token {TOKEN}"

def api_get(url):
"""统一入口:api.github.com / raw.githubusercontent.com 都带同一个头"""
req = urllib.request.Request(url, headers=HEADERS)
with urllib.request.urlopen(req, timeout=20) as r:
# 观测剩余额度(仅 api.github.com 返回,raw 无此头也安全)
print("remaining:", r.headers.get("X-RateLimit-Remaining"),
"/", r.headers.get("X-RateLimit-Limit"))
return r.read().decode("utf-8", "replace")

# 用法示例
forks = api_get("https://api.github.com/repos/eooce/choreo-2go/forks?per_page=100&page=1")
js = api_get("https://raw.githubusercontent.com/eooce/choreo-2go/main/files/index.js")





raw.githubusercontent.com 不设 REST 限流头,但同样接受 Authorization;
api.github.com 匿名 60 次/小时,带 Token 5000 次/小时;正规做法是环境变量传 token,别硬编码。

# 方式一:只对当前窗口生效(推荐,安全)
$env:GITHUB_TOKEN = "ghp_xxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxx"

# 方式二:永久写入用户环境变量(重开窗口也生效)
[Environment]::SetEnvironmentVariable("GITHUB_TOKEN", "ghp_xxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxx", "User")




脚本里只要读 os.environ.get("GITHUB_TOKEN"),有就用 Authorization: token $TOKEN(或 Bearer $TOKEN)请求头,没有就退回匿名——对所有 api.github.com 请求都管用,包括:

  • fetch_direct_forks.py / fetch_network_members.py / fetch_sub_forks.py(fork 枚举)
  • check_argo.py / check_indirect_forks.py / fetch_uuid.py(抓 files/index.js、index.js 原始文件)

13. 复现顺序(速查,2026-09-21 重建验证版)

⚠️ 跑下面 ①~⑭ 之前:

  1. 先按 §12.4 设置 $env:GITHUB_TOKEN,否则匿名 60 次/小时很容易在中途撞限流。
  2. PowerShell 窗口先执行编码初始化,避免中文输出破坏管道:
[Console]::OutputEncoding=[System.Text.Encoding]::UTF8
$env:PYTHONIOENCODING="utf-8"
python -X utf8 <脚本>.py



python fetch_direct_forks.py        # ① 直接 fork(582)→ forks_direct_full.json / direct_forks.txt
python fetch_network_members.py # ② network 页面 + 成员提取 → net_dump.html / network_members.txt
python analyze_net.py # ③(可选)结构分析
python extract_network.py # ④ repo key → net_repo_keys.json / indirect_forks_clean.txt(736)
python check_argo.py # ⑤ 93 仓库初轮 ARGO → argo_results.json
python check_all_forks.py # ⑥ 全部 fork ARGO → argo_all_results.json + custom_domains.txt(287) + file_content_cache.json
python check_indirect_forks.py # ⑦ 间接 fork ARGO → argo_indirect_results.json
python fetch_sub_forks.py # ⑧ BFS 子 fork(795,可断点续传)→ sub_forks_full.json + forks_all.json(796)
python extend_indirect.py # ⑧+ 子 fork 补检(BFS 新发现且 ⑦ 未覆盖的)并入 argo_indirect_results.json(796)
python check_bad_request.py # ⑨ 直接域名 Bad Request 探测(287 域名 → 17 bad)→ bad_request_results.json
python probe_indirect_domains.py # ⑩ 间接域名分类探测(332 域名 → 33 bad)→ probe_indirect_results.json
python fetch_uuid.py # ⑪ UUID 提取(351/351)→ uuid_results.json
python merge_final.py # ⑫ 收尾合并 → bad_request_all.json(33) + custom_domains.txt(332)
python build_relations.py # ⑬ 构建 repo↔domain↔uuid 关系(351/332)→ 5 个关系文件
python gen_vless.py # ⑭ 生成 VLESS 链接(351 条)→ vless_links.txt



时长参考(带 Token):①④ 约 5 分钟;⑤⑦ 约 25 分钟(内容 API 0.6s/次 × 并发 12);
⑧ BFS 约 15 分钟;⑨⑩ 约 5 分钟;⑪ 缓存命中近乎瞬时;⑫~⑭ 秒级。合计约 1 小时。
⑧ 中途断掉直接重跑即可(断点续传);⑤ 可跳过(小样本快照,结果不参与下游)。

13.5 完整脚本清单(15 个,全部内嵌于 §14)

# 脚本 作用
① fetch_direct_forks.py 枚举直接 fork(GitHub REST API 分页)
② fetch_network_members.py 抓取 Network members 全部分页 HTML
③ analyze_net.py 分析 network 页面结构(可选)
④ extract_network.py 从 HTML 清洗提取 repo key
⑤ check_argo.py 93 仓库初轮 ARGO 检测
⑥ check_all_forks.py 全部 fork ARGO 检测 + custom_domains.txt + 内容缓存
⑦ check_indirect_forks.py 间接 fork ARGO 检测
⑧ fetch_sub_forks.py BFS 枚举子 fork(断点续传)+ 收尾并集 forks_all.json
⑧+ extend_indirect.py 子 fork 补检(并入 argo_indirect_results.json)
⑨ check_bad_request.py 直接域名 Bad Request 探测
⑩ probe_indirect_domains.py 间接域名多状态分类探测
⑪ fetch_uuid.py UUID 提取(缓存优先)
⑫ merge_final.py 收尾合并(离线,无网络请求)
⑬ build_relations.py 构建 repo↔domain↔uuid 关系
⑭ gen_vless.py 生成 VLESS 链接

14. 附录:完整脚本归档(2026-09-21 重建验证版,可直接复制运行)

以下 15 个脚本即最新一次全量重建实际使用的版本(582 直接 fork → 796 全量 → 332 域名 → 351 UUID → 351 条 VLESS)。
复现方式:在空目录按 §13.5 顺序保存并执行即可。所有脚本自动读 GITHUB_TOKEN 环境变量;
统一 utf-8-sig 读取历史 JSON、sys.stdout.reconfigure 规避 PowerShell 中文坑;
⑥⑦⑪ 走 api.github.com contents API(绕开 raw CDN 卡顿)。

14.1 ① fetch_direct_forks.py

# -*- coding: utf-8 -*-
"""① 枚举 eooce/choreo-2go 直接 fork(GitHub REST API 分页,单页失败重试 5 次带退避)
输出: forks_direct_full.json / forks.json / forks_all.json / direct_forks.txt / all_forks.txt
"""
import os, json, time, sys, urllib.request, urllib.error

try:
sys.stdout.reconfigure(encoding="utf-8", errors="replace")
except Exception:
pass

ROOT = os.path.dirname(os.path.abspath(__file__))
OWNER_REPO = "eooce/choreo-2go"
TOKEN = os.environ.get("GITHUB_TOKEN", "")
HDRS = {"User-Agent": "Mozilla/5.0 (CodeArtsWork collect)",
"Accept": "application/vnd.github+json"}
if TOKEN:
HDRS["Authorization"] = f"token {TOKEN}"


def api_get(url):
req = urllib.request.Request(url, headers=HDRS)
with urllib.request.urlopen(req, timeout=30) as r:
rem = r.headers.get("X-RateLimit-Remaining")
if rem:
print(f" [rate] remaining={rem}/{r.headers.get('X-RateLimit-Limit')}")
return r.read().decode("utf-8", "replace")


def fetch_all_forks():
forks, page = [], 1
while True:
url = f"https://api.github.com/repos/{OWNER_REPO}/forks?per_page=100&page={page}"
data = None
for attempt in range(1, 6):
try:
data = json.loads(api_get(url))
break
except urllib.error.HTTPError as e:
if e.code == 403:
reset = e.headers.get("X-RateLimit-Reset")
wait = max(30, int(reset) - int(time.time()) + 5) if reset else 120
print(f" [403 rate-limit] sleep {min(wait, 600)}s (attempt {attempt}/5)")
time.sleep(min(wait, 600))
else:
print(f" [http {e.code}] retry {attempt}/5 ...")
time.sleep(5 * attempt)
except Exception as e:
print(f" [err] {e!r} retry {attempt}/5 ...")
time.sleep(5 * attempt)
if not data:
raise RuntimeError(f"page {page} failed after retries")
forks.extend(data)
print(f" page {page}: +{len(data)} (total {len(forks)})")
if len(data) < 100:
break
page += 1
return forks


def trim(f):
return {"full_name": f.get("full_name"),
"fork": f.get("fork"),
"forks_count": f.get("forks_count"),
"created_at": f.get("created_at"),
"pushed_at": f.get("pushed_at"),
"default_branch": f.get("default_branch")}


def main():
print(f"== (1) 枚举直接 fork: {OWNER_REPO} ==")
forks = fetch_all_forks()
trimmed = [trim(f) for f in forks]
names = sorted({t["full_name"] for t in trimmed if t["full_name"]})

def dump(fn, obj):
with open(os.path.join(ROOT, fn), "w", encoding="utf-8") as f:
json.dump(obj, f, ensure_ascii=False, indent=2)
print(f" wrote {fn}")

dump("forks_direct_full.json", trimmed)
dump("forks.json", trimmed)
dump("forks_all.json", trimmed)
with open(os.path.join(ROOT, "direct_forks.txt"), "w", encoding="utf-8") as f:
f.write("\n".join(names) + "\n")
with open(os.path.join(ROOT, "all_forks.txt"), "w", encoding="utf-8") as f:
f.write("\n".join(names) + "\n")
total_sub = sum(t["forks_count"] or 0 for t in trimmed)
print(f"直接 fork 总数: {len(names)};子 fork 数合计: {total_sub}")


if __name__ == "__main__":
main()



14.2 ② fetch_network_members.py

# -*- coding: utf-8 -*-
"""② 抓取 network/members 全部分页 HTML,合并保存为 net_dump.html,
并提取 repo 链接(network_members.txt)与 hovercard owner(network_members_owners.txt)"""
import os, re, sys, time, urllib.request

try:
sys.stdout.reconfigure(encoding="utf-8", errors="replace")
except Exception:
pass

ROOT = os.path.dirname(os.path.abspath(__file__))
ROOT_REPO = "eooce/choreo-2go"
TOKEN = os.environ.get("GITHUB_TOKEN", "")
HDRS = {"User-Agent": "Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 "
"(KHTML, like Gecko) Chrome/124.0 Safari/537.36",
"Accept": "text/html,application/xhtml+xml"}
if TOKEN:
HDRS["Authorization"] = f"token {TOKEN}"

RE_REPO = re.compile(r'href="/([^/"]+/[^/"]+)"')
RE_HOVER = re.compile(r'data-hovercard-url="/users/([^/"]+)/hovercard"')


def fetch(url):
for a in range(1, 6):
try:
req = urllib.request.Request(url, headers=HDRS)
with urllib.request.urlopen(req, timeout=30) as r:
return r.read().decode("utf-8", "replace")
except Exception as e:
print(f" [retry {a}/5] {e!r}")
time.sleep(3 * a)
raise RuntimeError("fetch failed: " + url)


def main():
print(f"== (2) network members: {ROOT_REPO} ==")
all_repos, all_owners = set(), set()
page, stagnant = 1, 0
dump_parts = []
while page <= 60 and stagnant < 2:
base = f"https://github.com/{ROOT_REPO}/network/members"
url = base if page == 1 else f"{base}?page={page}"
html = fetch(url)
dump_parts.append(f"<!-- ===== page {page} ===== -->\n{html}")
repos = {m.group(1).strip("/") for m in RE_REPO.finditer(html)}
owners = set(RE_HOVER.findall(html))
new = repos - all_repos
print(f" page {page}: repos={len(repos)} new={len(new)} owners_total={len(all_owners | owners)}")
all_repos |= repos
all_owners |= owners
stagnant = stagnant + 1 if not new else 0
page += 1
time.sleep(1)

with open(os.path.join(ROOT, "net_dump.html"), "w", encoding="utf-8") as f:
f.write("\n".join(dump_parts))
with open(os.path.join(ROOT, "network_members.txt"), "w", encoding="utf-8") as f:
f.write("\n".join(sorted(all_repos)) + "\n")
with open(os.path.join(ROOT, "network_members_owners.txt"), "w", encoding="utf-8") as f:
f.write("\n".join(sorted(all_owners)) + "\n")
print(f"network members 粗提取: {len(all_repos)} repos / {len(all_owners)} owners(含干扰,稍后由 extract_network 清洗)")


if __name__ == "__main__":
main()



14.3 ③ analyze_net.py

# -*- coding: utf-8 -*-
"""③ 分析 net_dump.html 结构(可选步骤)"""
import os, re, sys

try:
sys.stdout.reconfigure(encoding="utf-8", errors="replace")
except Exception:
pass

ROOT = os.path.dirname(os.path.abspath(__file__))
html = open(os.path.join(ROOT, "net_dump.html"), encoding="utf-8", errors="replace").read()

print("size:", len(html))
print("repo hrefs:", len(re.findall(r'href="/[^/"]+/[^/"]+"', html)))
print("hovercards:", len(re.findall(r'data-hovercard-url="/users/', html)))

pages = sorted(set(re.findall(r'href="([^"]*network/members[^"]*)"', html)))
print("page links:", pages if len(pages) < 20 else f"{len(pages)} links (sample {pages[:5]})")

hover = re.findall(r'data-hovercard-url="([^"]+)"', html)
print("hovercard sample:", hover[:5])

m = re.search(r'aria-label="Pagination".{0,4000}?</nav>', html, re.S)
print("pagination block:", "found" if m else "not found")



14.4 ④ extract_network.py

# -*- coding: utf-8 -*-
"""④ 从 net_dump.html 提取 repo key
输出: net_repo_keys.json / indirect_forks.txt / indirect_forks_clean.txt
(清洗:排除 root 仓库、非 owner/repo 形态、导航/静态资源干扰链接)"""
import os, re, json, sys

try:
sys.stdout.reconfigure(encoding="utf-8", errors="replace")
except Exception:
pass

ROOT = os.path.dirname(os.path.abspath(__file__))
ROOT_REPO = "eooce/choreo-2go"

html = open(os.path.join(ROOT, "net_dump.html"), encoding="utf-8", errors="replace").read()
raw = set(re.findall(r'href="/([^/"]+/[^/"]+)"', html))

BANNED_PREFIX = ("features/", "settings/", "topics/", "collections/", "notifications",
"login", "join", "explore", "apps/", "orgs/", "account/", "marketplace",
"pulls", "issues", "search", "security", "pricing", "customer-stories",
"readme", "about", "site/", "enterprise", "team", "sponsors", "new",
"codespaces", "generative-ai", "contact", "terms", "privacy")
STATIC_EXT = (".js", ".json", ".md", ".py", ".html", ".css", ".png", ".svg", ".txt")


def ok(r):
if not re.fullmatch(r"[\w.\-]+/[\w.\-]+", r):
return False
if r.lower().endswith(STATIC_EXT):
return False
if any(r.startswith(b) or r.lower().startswith(b) for b in BANNED_PREFIX):
return False
return True


keys = sorted({r for r in raw if ok(r) and r.lower() != ROOT_REPO.lower()})

with open(os.path.join(ROOT, "net_repo_keys.json"), "w", encoding="utf-8") as f:
json.dump(keys, f, ensure_ascii=False, indent=2)
with open(os.path.join(ROOT, "indirect_forks.txt"), "w", encoding="utf-8") as f:
f.write("\n".join(keys) + "\n")
with open(os.path.join(ROOT, "indirect_forks_clean.txt"), "w", encoding="utf-8") as f:
f.write("\n".join(keys) + "\n")

print(f"原始 href 候选: {len(raw)};清洗后间接 fork: {len(keys)}")



14.5 ⑤ check_argo.py

# -*- coding: utf-8 -*-
"""⑤ 初轮 ARGO 域名提取(93 仓库)
原始版本硬编码 93 个 owner/repo;重建版改为取 direct_forks.txt 排序后前 93 个(确定性等价)。
输出: argo_results.json
"""
import os, re, json, sys, time, urllib.request, urllib.error, concurrent.futures

try:
sys.stdout.reconfigure(encoding="utf-8", errors="replace")
except Exception:
pass

ROOT = os.path.dirname(os.path.abspath(__file__))
ARGO_RE = re.compile(r"ARGO_DOMAIN\s*=\s*process\.env\.ARGO_DOMAIN\s*\|\|\s*(['\"])(.*?)\1")
HDRS = {"User-Agent": "Mozilla/5.0 (CodeArtsWork check)",
"Accept": "application/vnd.github.raw"}
if os.environ.get("GITHUB_TOKEN"):
HDRS["Authorization"] = f"token {os.environ['GITHUB_TOKEN']}"


def get_file(repo, path):
"""api.github.com contents API(raw CDN 在本环境卡顿),main→master 回退"""
for branch in ("main", "master"):
url = f"https://api.github.com/repos/{repo}/contents/{path}?ref={branch}"
for attempt in range(3):
try:
with urllib.request.urlopen(urllib.request.Request(url, headers=HDRS), timeout=20) as r:
return ("ok", r.read().decode("utf-8", "replace"))
except urllib.error.HTTPError as e:
if e.code == 404:
break
if e.code in (403, 429):
time.sleep(60)
continue
time.sleep(2 * attempt)
except Exception:
time.sleep(2 * attempt)
return ("404", "")


def check_repo(repo):
res = {"repo": repo}
st1, c1 = get_file(repo, "files/index.js") # 首选 files/index.js
m1 = ARGO_RE.search(c1) if st1 == "ok" else None
res["files_idx_fallback"] = m1.group(2) if m1 else None
st2, c2 = get_file(repo, "index.js") # 次选根 index.js
m2 = ARGO_RE.search(c2) if st2 == "ok" else None
res["root_idx_fallback"] = m2.group(2) if m2 else None
return res


def main():
repos = [l.strip() for l in open(os.path.join(ROOT, "direct_forks.txt"),
encoding="utf-8-sig") if l.strip()]
repos = sorted(repos)[:93]
print(f"== (5) 初轮 ARGO 检测: {len(repos)} 仓库 ==")
out = []
with concurrent.futures.ThreadPoolExecutor(max_workers=12) as ex:
futs = {ex.submit(check_repo, r): r for r in repos}
for i, fut in enumerate(concurrent.futures.as_completed(futs), 1):
out.append(fut.result())
if i % 20 == 0:
print(f" {i}/{len(repos)}")
out.sort(key=lambda r: r["repo"])
with open(os.path.join(ROOT, "argo_results.json"), "w", encoding="utf-8") as f:
json.dump(out, f, ensure_ascii=False, indent=2)
ff = [r for r in out if r["files_idx_fallback"]]
rf = [r for r in out if r["root_idx_fallback"]]
nf = [r for r in out if not r["files_idx_fallback"] and not r["root_idx_fallback"]]
print(f"files/index.js 非空 fallback: {len(ff)};根 index.js 非空: {len(rf)};空/缺失: {len(nf)}")


if __name__ == "__main__":
main()



14.6 ⑥ check_all_forks.py

# -*- coding: utf-8 -*-
"""⑥ 全部 fork 的 ARGO 域名提取 → argo_all_results.json
附带:汇总直接 fork 的自定义域名 → custom_domains.txt(junk 过滤 + 剔除默认域名)。
文件抓取走 api.github.com contents API(raw.githubusercontent.com 在本环境严重卡顿);
抓到的 index.js 内容写入 file_content_cache.json 供 ⑪ fetch_uuid 复用(省 API 配额)。
注意:所有读取方统一用 encoding="utf-8-sig"(历史版本首条记录带 BOM)。
"""
import os, re, json, sys, time, urllib.request, urllib.error, concurrent.futures

try:
sys.stdout.reconfigure(encoding="utf-8", errors="replace")
except Exception:
pass

ROOT = os.path.dirname(os.path.abspath(__file__))
ARGO_RE = re.compile(r"ARGO_DOMAIN\s*=\s*process\.env\.ARGO_DOMAIN\s*\|\|\s*(['\"])(.*?)\1")
JUNK = {"", "NOT_FOUND", "NO_MATCH", "OBFUSCATED", "error"}
DEFAULT_DOMAIN = "choreo.zzx.free.hr"
TOKEN = os.environ.get("GITHUB_TOKEN", "")
HDRS = {"User-Agent": "Mozilla/5.0 (CodeArtsWork check)",
"Accept": "application/vnd.github.raw"}
if TOKEN:
HDRS["Authorization"] = f"token {TOKEN}"


class RateLimited(Exception):
pass


def get_file(repo, path):
"""api.github.com contents API,main→master 回退,重试 3 次"""
for branch in ("main", "master"):
url = f"https://api.github.com/repos/{repo}/contents/{path}?ref={branch}"
for attempt in range(3):
try:
with urllib.request.urlopen(urllib.request.Request(url, headers=HDRS), timeout=20) as r:
rem = r.headers.get("X-RateLimit-Remaining")
if rem == "0":
raise RateLimited("remaining=0")
return ("ok", r.read().decode("utf-8", "replace"))
except urllib.error.HTTPError as e:
if e.code == 404:
break # 下一 branch
if e.code in (403, 429):
reset = e.headers.get("X-RateLimit-Reset")
wait = min(600, max(20, int(reset) - int(time.time()) + 5)) if reset else 120
print(f" [rate-limit] sleep {wait}s")
time.sleep(wait)
continue
time.sleep(2 * attempt)
except RateLimited:
reset = None
try:
with urllib.request.urlopen(urllib.request.Request(
"https://api.github.com/rate_limit", headers=HDRS), timeout=15) as r:
reset = json.loads(r.read())["resources"]["core"]["reset"]
except Exception:
pass
wait = min(600, max(20, int(reset) - int(time.time()) + 5)) if reset else 120
print(f" [rate-limit remaining=0] sleep {wait}s")
time.sleep(wait)
except Exception:
time.sleep(2 * attempt)
return ("404", "")


def check_repo(repo):
res = {"repo": repo}
st1, c1 = get_file(repo, "files/index.js")
m1 = ARGO_RE.search(c1) if st1 == "ok" else None
res["files_idx_fallback"] = m1.group(2) if m1 else None
st2, c2 = get_file(repo, "index.js")
m2 = ARGO_RE.search(c2) if st2 == "ok" else None
res["root_idx_fallback"] = m2.group(2) if m2 else None
res["_cache"] = {}
if st1 == "ok":
res["_cache"]["files/index.js"] = c1
if st2 == "ok":
res["_cache"]["index.js"] = c2
return res


def valid_domain(d):
if not d:
return False
d = d.strip().rstrip("/").strip()
if d in JUNK or d == DEFAULT_DOMAIN:
return False
if "ip6.arpa" in d or d == "free.hr" or d.endswith(".free.hr") or d == "cho.zzx.free":
return False
return True


def main():
data = json.load(open(os.path.join(ROOT, "forks_all.json"), encoding="utf-8-sig"))
repos = sorted({d["full_name"] for d in data if d.get("full_name")})
print(f"== (6) 全部 fork ARGO 检测: {len(repos)} 仓库 ==")
out = []
with concurrent.futures.ThreadPoolExecutor(max_workers=12) as ex:
futs = {ex.submit(check_repo, r): r for r in repos}
for i, fut in enumerate(concurrent.futures.as_completed(futs), 1):
out.append(fut.result())
if i % 50 == 0:
print(f" {i}/{len(repos)}")
out.sort(key=lambda r: r["repo"])
# 拆缓存:argo_all_results 只留元数据,文件内容进 file_content_cache.json
cache = {}
for r in out:
c = r.pop("_cache", {})
if c:
cache[r["repo"]] = c
with open(os.path.join(ROOT, "argo_all_results.json"), "w", encoding="utf-8") as f:
json.dump(out, f, ensure_ascii=False, indent=2)
with open(os.path.join(ROOT, "file_content_cache.json"), "w", encoding="utf-8") as f:
json.dump(cache, f, ensure_ascii=False)

doms = set()
for r in out:
for k in ("files_idx_fallback", "root_idx_fallback"):
v = r.get(k)
if valid_domain(v):
doms.add(v.strip().rstrip("/").strip())
with open(os.path.join(ROOT, "custom_domains.txt"), "w", encoding="utf-8") as f:
f.write("\n".join(sorted(doms)) + "\n")

ff = sum(1 for r in out if r["files_idx_fallback"])
rf = sum(1 for r in out if r["root_idx_fallback"])
print(f"完成:files 非空 {ff} / root 非空 {rf} / custom_domains.txt {len(doms)} 域名")


if __name__ == "__main__":
main()



14.7 ⑦ check_indirect_forks.py

# -*- coding: utf-8 -*-
"""⑦ 间接 fork(network members)ARGO 检测 → argo_indirect_results.json
每条记录 files_index / root_index 两个 fallback 候选。
文件抓取走 api.github.com contents API;内容并入 file_content_cache.json 供 ⑪ 复用。
输入: indirect_forks_clean.txt(extract_network.py 产出),并与 net_repo_keys.json 取并集。
"""
import os, re, json, sys, time, urllib.request, urllib.error, concurrent.futures

try:
sys.stdout.reconfigure(encoding="utf-8", errors="replace")
except Exception:
pass

ROOT = os.path.dirname(os.path.abspath(__file__))
ARGO_RE = re.compile(r"ARGO_DOMAIN\s*=\s*process\.env\.ARGO_DOMAIN\s*\|\|\s*(['\"])(.*?)\1")
ROOT_REPO = "eooce/choreo-2go"
TOKEN = os.environ.get("GITHUB_TOKEN", "")
HDRS = {"User-Agent": "Mozilla/5.0 (CodeArtsWork check)",
"Accept": "application/vnd.github.raw"}
if TOKEN:
HDRS["Authorization"] = f"token {TOKEN}"


def get_file(repo, path):
"""api.github.com contents API,main→master 回退,重试 3 次"""
for branch in ("main", "master"):
url = f"https://api.github.com/repos/{repo}/contents/{path}?ref={branch}"
for attempt in range(3):
try:
with urllib.request.urlopen(urllib.request.Request(url, headers=HDRS), timeout=20) as r:
rem = r.headers.get("X-RateLimit-Remaining")
if rem == "0":
raise RateLimited("remaining=0")
return ("ok", r.read().decode("utf-8", "replace"))
except urllib.error.HTTPError as e:
if e.code == 404:
break
if e.code in (403, 429):
reset = e.headers.get("X-RateLimit-Reset")
wait = min(600, max(20, int(reset) - int(time.time()) + 5)) if reset else 120
print(f" [rate-limit] sleep {wait}s")
time.sleep(wait)
continue
time.sleep(2 * attempt)
except RateLimited:
reset = None
try:
with urllib.request.urlopen(urllib.request.Request(
"https://api.github.com/rate_limit", headers=HDRS), timeout=15) as r:
reset = json.loads(r.read())["resources"]["core"]["reset"]
except Exception:
pass
wait = min(600, max(20, int(reset) - int(time.time()) + 5)) if reset else 120
print(f" [rate-limit remaining=0] sleep {wait}s")
time.sleep(wait)
except Exception:
time.sleep(2 * attempt)
return ("404", "")


class RateLimited(Exception):
pass


def check_repo(repo):
res = {"repo": repo}
st1, c1 = get_file(repo, "files/index.js")
m1 = ARGO_RE.search(c1) if st1 == "ok" else None
res["files_index"] = m1.group(2) if m1 else None
st2, c2 = get_file(repo, "index.js")
m2 = ARGO_RE.search(c2) if st2 == "ok" else None
res["root_index"] = m2.group(2) if m2 else None
res["_cache"] = {}
if st1 == "ok":
res["_cache"]["files/index.js"] = c1
if st2 == "ok":
res["_cache"]["index.js"] = c2
return res


def main():
repos = set()
p1 = os.path.join(ROOT, "indirect_forks_clean.txt")
p2 = os.path.join(ROOT, "net_repo_keys.json")
if os.path.exists(p1):
repos |= {l.strip() for l in open(p1, encoding="utf-8-sig") if l.strip()}
if os.path.exists(p2):
repos |= set(json.load(open(p2, encoding="utf-8-sig")))
repos = sorted(r for r in repos if r and r.lower() != ROOT_REPO.lower())
print(f"== (7) 间接 fork ARGO 检测: {len(repos)} 仓库 ==")
out = []
with concurrent.futures.ThreadPoolExecutor(max_workers=12) as ex:
futs = {ex.submit(check_repo, r): r for r in repos}
for i, fut in enumerate(concurrent.futures.as_completed(futs), 1):
out.append(fut.result())
if i % 50 == 0:
print(f" {i}/{len(repos)}")
out.sort(key=lambda r: r["repo"])
cache_path = os.path.join(ROOT, "file_content_cache.json")
cache = json.load(open(cache_path, encoding="utf-8-sig")) if os.path.exists(cache_path) else {}
for r in out:
c = r.pop("_cache", {})
if c:
cache.setdefault(r["repo"], {}).update(c)
with open(os.path.join(ROOT, "argo_indirect_results.json"), "w", encoding="utf-8") as f:
json.dump(out, f, ensure_ascii=False, indent=2)
with open(cache_path, "w", encoding="utf-8") as f:
json.dump(cache, f, ensure_ascii=False)
fi = sum(1 for r in out if r["files_index"])
ri = sum(1 for r in out if r["root_index"])
print(f"完成:files_index 非空 {fi} / root_index 非空 {ri}")


if __name__ == "__main__":
main()



14.8 ⑧ fetch_sub_forks.py

# -*- coding: utf-8 -*-
"""⑧ BFS 枚举子 fork(间接 fork 的 fork),带断点续传
状态: sub_forks_state.json(visited / queue / 累积结果索引)
结果: sub_forks_full.json
限流处理: X-RateLimit-Remaining=0 或 403 时把当前任务重新入队、保存状态、优雅退出。
收尾:forks_all.json = 直接 fork + 间接 fork + 子 fork 并集(trimmed 全量)。
"""
import os, json, time, sys, urllib.request, urllib.error
from collections import deque

try:
sys.stdout.reconfigure(encoding="utf-8", errors="replace")
except Exception:
pass

ROOT = os.path.dirname(os.path.abspath(__file__))
HDRS = {"User-Agent": "Mozilla/5.0 (CodeArtsWork collect)",
"Accept": "application/vnd.github+json"}
if os.environ.get("GITHUB_TOKEN"):
HDRS["Authorization"] = f"token {os.environ['GITHUB_TOKEN']}"


class RateLimited(Exception):
pass


def api_get(url):
req = urllib.request.Request(url, headers=HDRS)
try:
with urllib.request.urlopen(req, timeout=30) as r:
rem = r.headers.get("X-RateLimit-Remaining")
body = r.read().decode("utf-8", "replace")
if rem == "0":
raise RateLimited("remaining=0")
return json.loads(body)
except urllib.error.HTTPError as e:
if e.code in (403, 429):
raise RateLimited(f"http{e.code}")
if e.code == 404:
return None
raise


def load_state():
sp = os.path.join(ROOT, "sub_forks_state.json")
if os.path.exists(sp):
return json.load(open(sp, encoding="utf-8-sig"))
seeds = []
p = os.path.join(ROOT, "indirect_forks_clean.txt")
if os.path.exists(p):
seeds = [l.strip() for l in open(p, encoding="utf-8-sig") if l.strip()]
return {"visited": [], "queue": seeds, "results": []}


def save_state(st):
with open(os.path.join(ROOT, "sub_forks_state.json"), "w", encoding="utf-8") as f:
json.dump(st, f, ensure_ascii=False)
with open(os.path.join(ROOT, "sub_forks_full.json"), "w", encoding="utf-8") as f:
json.dump(st["results"], f, ensure_ascii=False, indent=2)


def fetch_children(repo):
children = []
page = 1
while True:
data = api_get(f"https://api.github.com/repos/{repo}/forks?per_page=100&page={page}")
if not data:
break
children.extend(data)
if len(data) < 100:
break
page += 1
return children


def main():
st = load_state()
visited = set(st["visited"])
queue = deque(r for r in st["queue"] if r not in visited)
results = st["results"]
have = {r["full_name"] for r in results}
print(f"== (8) BFS 子 fork: 已访问 {len(visited)},队列 {len(queue)},已有结果 {len(results)} ==")

count = 0
while queue:
repo = queue.popleft()
if repo in visited:
continue
try:
children = fetch_children(repo)
except RateLimited as e:
queue.appendleft(repo)
st.update(visited=sorted(visited), queue=list(queue), results=results)
save_state(st)
print(f"[rate-limit {e}] 任务重新入队,状态已保存,退出。剩余队列 {len(queue)}")
return
except Exception as e:
print(f" [err] {repo}: {e!r}(跳过)")
visited.add(repo)
continue

visited.add(repo)
for c in children:
fn = c.get("full_name")
if not fn or fn in have or fn == repo:
continue
have.add(fn)
results.append({"full_name": fn, "parent": repo,
"forks_count": c.get("forks_count"),
"created_at": c.get("created_at"),
"pushed_at": c.get("pushed_at"),
"default_branch": c.get("default_branch")})
queue.append(fn)
count += 1
if count % 25 == 0:
st.update(visited=sorted(visited), queue=list(queue), results=results)
save_state(st)
print(f" visited={len(visited)} queue={len(queue)} results={len(results)}")
time.sleep(0.15)

st.update(visited=sorted(visited), queue=[], results=results)
save_state(st)
print(f"BFS 完成:访问 {len(visited)},子 fork 共 {len(results)}")

# 收尾:forks_all.json = 直接 + 间接 + 子 fork 并集
allnames = set()
p = os.path.join(ROOT, "forks_direct_full.json")
if os.path.exists(p):
allnames |= {d["full_name"] for d in json.load(open(p, encoding="utf-8-sig")) if d.get("full_name")}
allnames |= {r["full_name"] for r in results}
p = os.path.join(ROOT, "indirect_forks_clean.txt")
if os.path.exists(p):
allnames |= {l.strip() for l in open(p, encoding="utf-8-sig") if l.strip()}
merged = [{"full_name": n, "source": "subfork_union"} for n in sorted(allnames)]
with open(os.path.join(ROOT, "forks_all.json"), "w", encoding="utf-8") as f:
json.dump(merged, f, ensure_ascii=False, indent=2)
with open(os.path.join(ROOT, "all_forks.txt"), "w", encoding="utf-8") as f:
f.write("\n".join(sorted(allnames)) + "\n")
print(f"forks_all.json / all_forks.txt: {len(allnames)} 全量 fork")


if __name__ == "__main__":
main()



14.9 ⑧+ extend_indirect.py

# -*- coding: utf-8 -*-
"""⑧+ 子 fork 补检:对 BFS 发现但 ⑦ 未覆盖的仓库做 ARGO 检测,
结果并入 argo_indirect_results.json + file_content_cache.json"""
import os, re, json, sys, time, urllib.request, urllib.error, concurrent.futures

try:
sys.stdout.reconfigure(encoding="utf-8", errors="replace")
except Exception:
pass

ROOT = os.path.dirname(os.path.abspath(__file__))
ARGO_RE = re.compile(r"ARGO_DOMAIN\s*=\s*process\.env\.ARGO_DOMAIN\s*\|\|\s*(['\"])(.*?)\1")
HDRS = {"User-Agent": "Mozilla/5.0 (CodeArtsWork check)",
"Accept": "application/vnd.github.raw"}
if os.environ.get("GITHUB_TOKEN"):
HDRS["Authorization"] = f"token {os.environ['GITHUB_TOKEN']}"


def get_file(repo, path):
for branch in ("main", "master"):
url = f"https://api.github.com/repos/{repo}/contents/{path}?ref={branch}"
for attempt in range(3):
try:
with urllib.request.urlopen(urllib.request.Request(url, headers=HDRS), timeout=20) as r:
return ("ok", r.read().decode("utf-8", "replace"))
except urllib.error.HTTPError as e:
if e.code == 404:
break
if e.code in (403, 429):
time.sleep(60)
continue
time.sleep(2 * attempt)
except Exception:
time.sleep(2 * attempt)
return ("404", "")


def check_repo(repo):
res = {"repo": repo}
st1, c1 = get_file(repo, "files/index.js")
m1 = ARGO_RE.search(c1) if st1 == "ok" else None
res["files_index"] = m1.group(2) if m1 else None
st2, c2 = get_file(repo, "index.js")
m2 = ARGO_RE.search(c2) if st2 == "ok" else None
res["root_index"] = m2.group(2) if m2 else None
res["_cache"] = {}
if st1 == "ok":
res["_cache"]["files/index.js"] = c1
if st2 == "ok":
res["_cache"]["index.js"] = c2
return res


def main():
sub = json.load(open(os.path.join(ROOT, "sub_forks_full.json"), encoding="utf-8-sig"))
ind = json.load(open(os.path.join(ROOT, "argo_indirect_results.json"), encoding="utf-8-sig"))
covered = {r["repo"] for r in ind}
todo = sorted({r["full_name"] for r in sub} - covered)
print(f"子 fork 补检: {len(todo)} 仓库")
if not todo:
return
new = []
with concurrent.futures.ThreadPoolExecutor(max_workers=12) as ex:
futs = {ex.submit(check_repo, r): r for r in todo}
for i, fut in enumerate(concurrent.futures.as_completed(futs), 1):
new.append(fut.result())
if i % 20 == 0:
print(f" {i}/{len(todo)}")
cache_path = os.path.join(ROOT, "file_content_cache.json")
cache = json.load(open(cache_path, encoding="utf-8-sig")) if os.path.exists(cache_path) else {}
for r in new:
c = r.pop("_cache", {})
if c:
cache.setdefault(r["repo"], {}).update(c)
merged = sorted(ind + new, key=lambda r: r["repo"])
with open(os.path.join(ROOT, "argo_indirect_results.json"), "w", encoding="utf-8") as f:
json.dump(merged, f, ensure_ascii=False, indent=2)
with open(cache_path, "w", encoding="utf-8") as f:
json.dump(cache, f, ensure_ascii=False)
fi = sum(1 for r in new if r["files_index"])
ri = sum(1 for r in new if r["root_index"])
print(f"补检完成:新增 {len(new)}(files 非空 {fi} / root 非空 {ri}),"
f"argo_indirect 总数 {len(merged)},缓存 {len(cache)}")


if __name__ == "__main__":
main()



14.10 ⑨ check_bad_request.py

# -*- coding: utf-8 -*-
"""⑨ 直接域名 Bad Request 探测 → bad_request_results.json
输入: custom_domains.txt
逻辑: 每个域名先 https:// 后 http://,verify=False、allow_redirects=True;
判定 bad_request: status_code == 400 或响应体包含 "Bad Request";其余归 other。
输出: bad_request_results.json {"bad_request": [...], "other": [...]}
"""
import os, json, sys, warnings, concurrent.futures
import requests
import urllib3

try:
sys.stdout.reconfigure(encoding="utf-8", errors="replace")
except Exception:
pass
urllib3.disable_warnings()

ROOT = os.path.dirname(os.path.abspath(__file__))
HEADERS = {"User-Agent": "Mozilla/5.0 (CodeArtsWork probe)"}


def normal(d):
return d.strip().rstrip("/").strip()


def probe(domain):
res = {"domain": domain, "scheme": None, "status": None, "kind": None, "snippet": ""}
last = ""
for scheme in ("https", "http"):
try:
r = requests.get(f"{scheme}://{domain}", headers=HEADERS, timeout=15,
verify=False, allow_redirects=True)
text = r.text[:3000]
res.update(scheme=scheme, status=r.status_code)
res["kind"] = "bad_request" if (r.status_code == 400 or "Bad Request" in text) else "other"
res["snippet"] = text[:200]
return res
except Exception as e:
last = repr(e)[:200]
res["kind"] = "other"
res["snippet"] = last or "unreachable"
return res


def main():
domains = sorted({normal(l) for l in open(os.path.join(ROOT, "custom_domains.txt"),
encoding="utf-8-sig") if normal(l)})
print(f"== (9) Bad Request 探测: {len(domains)} 域名 ==")
results = []
with concurrent.futures.ThreadPoolExecutor(max_workers=10) as ex:
futs = {ex.submit(probe, d): d for d in domains}
for i, fut in enumerate(concurrent.futures.as_completed(futs), 1):
results.append(fut.result())
if i % 50 == 0:
print(f" {i}/{len(domains)}")
results.sort(key=lambda r: r["domain"])
bad = [r for r in results if r["kind"] == "bad_request"]
other = [r for r in results if r["kind"] != "bad_request"]
with open(os.path.join(ROOT, "bad_request_results.json"), "w", encoding="utf-8") as f:
json.dump({"bad_request": bad, "other": other}, f, ensure_ascii=False, indent=2)
print(f"bad_request: {len(bad)};other: {len(other)}")
for r in bad:
print(f" [400] {r['domain']} (scheme={r['scheme']}, status={r['status']})")


if __name__ == "__main__":
main()



14.11 ⑩ probe_indirect_domains.py

# -*- coding: utf-8 -*-
"""⑩ 间接域名探测(多状态分类)→ probe_indirect_results.json
输入: argo_indirect_results.json 中的有效 fallback 域名(junk 过滤)
分类: 400_BAD_REQUEST / 502_BACKEND_DOWN / 530_TUNNEL_DOWN / 530_OTHER / 200_OK /
DNS_FAIL / CONN_FAIL / ERROR / UNREACHABLE / HTTP_xxx
每条含 domain / scheme / status / kind / snippet / repos
"""
import os, re, json, sys, warnings, concurrent.futures
import requests
import urllib3

try:
sys.stdout.reconfigure(encoding="utf-8", errors="replace")
except Exception:
pass
urllib3.disable_warnings()

ROOT = os.path.dirname(os.path.abspath(__file__))
HEADERS = {"User-Agent": "Mozilla/5.0 (CodeArtsWork probe)"}
JUNK = {"", "NOT_FOUND", "NO_MATCH", "OBFUSCATED", "error"}
DEFAULT_DOMAIN = "choreo.zzx.free.hr"


def normal(d):
return (d or "").strip().rstrip("/").strip()


def collect_domains():
data = json.load(open(os.path.join(ROOT, "argo_indirect_results.json"), encoding="utf-8-sig"))
d2repos = {}
for r in data:
for k in ("files_index", "root_index"):
v = normal(r.get(k))
if not v or v in JUNK or v == DEFAULT_DOMAIN:
continue
if "ip6.arpa" in v or v == "free.hr" or v.endswith(".free.hr") or v == "cho.zzx.free":
continue
d2repos.setdefault(v, set()).add(r["repo"])
return {d: sorted(rs) for d, rs in sorted(d2repos.items())}


def classify(domain):
res = {"domain": domain, "scheme": None, "status": None, "kind": None, "snippet": ""}
last = ""
for scheme in ("https", "http"):
try:
r = requests.get(f"{scheme}://{domain}", headers=HEADERS, timeout=15,
verify=False, allow_redirects=True)
status, text = r.status_code, r.text[:3000]
res.update(scheme=scheme, status=status)
if status == 400 or "Bad Request" in text:
res["kind"] = "400_BAD_REQUEST"
elif status == 502:
res["kind"] = "502_BACKEND_DOWN"
elif status == 530 and "1033" in text:
res["kind"] = "530_TUNNEL_DOWN"
elif status == 530:
res["kind"] = "530_OTHER"
elif status == 200:
res["kind"] = "200_OK"
else:
res["kind"] = f"HTTP_{status}"
res["snippet"] = text[:300]
return res
except requests.exceptions.SSLError:
last = "ssl error"
continue
except requests.exceptions.ConnectionError as e:
s = repr(e)
if "getaddrinfo" in s or "Name or service not known" in s or "nodename nor servname" in s:
res["kind"] = "DNS_FAIL"
res["snippet"] = "getaddrinfo failed"
else:
res["kind"] = "CONN_FAIL"
res["snippet"] = s[:200]
res["scheme"] = scheme
return res
except Exception as e:
res.update(scheme=scheme, status="ERR", kind="ERROR", snippet=repr(e)[:200])
return res
res["kind"] = "UNREACHABLE"
res["snippet"] = last
return res


def main():
d2repos = collect_domains()
domains = list(d2repos)
print(f"== (10) 间接域名探测: {len(domains)} 唯一域名 ==")
results = []
with concurrent.futures.ThreadPoolExecutor(max_workers=10) as ex:
futs = {ex.submit(classify, d): d for d in domains}
for i, fut in enumerate(concurrent.futures.as_completed(futs), 1):
r = fut.result()
r["repos"] = d2repos[r["domain"]]
results.append(r)
if i % 20 == 0:
print(f" {i}/{len(domains)}")
results.sort(key=lambda r: r["domain"])
with open(os.path.join(ROOT, "probe_indirect_results.json"), "w", encoding="utf-8") as f:
json.dump(results, f, ensure_ascii=False, indent=2)
from collections import Counter
cnt = Counter(r["kind"] for r in results)
print("分类统计:", dict(cnt))
b400 = [r["domain"] for r in results if r["kind"] == "400_BAD_REQUEST"]
print(f"400_BAD_REQUEST: {len(b400)}")
for d in b400:
print(f" [400] {d}")


if __name__ == "__main__":
main()



14.12 ⑪ fetch_uuid.py

# -*- coding: utf-8 -*-
"""⑪ UUID 提取 → uuid_results.json
优先读 ⑥⑦ 产出的 file_content_cache.json(零 API 调用);
缓存缺失的仓库再走 api.github.com contents API 补抓。
候选 repo = argo_all_results ∪ argo_indirect_results 中存在有效自定义域名的 repo。
正则(通用化,允许引号与空白差异):
const UUID = process.env.UUID || '<36位>'
输出: {repo: uuid | "NOT_FOUND" | "error: ..."}
"""
import os, re, json, sys, time, urllib.request, urllib.error, concurrent.futures

try:
sys.stdout.reconfigure(encoding="utf-8", errors="replace")
except Exception:
pass

ROOT = os.path.dirname(os.path.abspath(__file__))
UUID_RE = re.compile(r"const\s+UUID\s*=\s*process\.env\.UUID\s*\|\|\s*['\"]([0-9a-fA-F-]{36})['\"]")
JUNK = {"", "NOT_FOUND", "NO_MATCH", "OBFUSCATED", "error"}
DEFAULT_DOMAIN = "choreo.zzx.free.hr"
HDRS = {"User-Agent": "Mozilla/5.0 (CodeArtsWork uuid)",
"Accept": "application/vnd.github.raw"}
if os.environ.get("GITHUB_TOKEN"):
HDRS["Authorization"] = f"token {os.environ['GITHUB_TOKEN']}"


def get_file(repo, path):
"""api.github.com contents API,main→master 回退"""
for branch in ("main", "master"):
url = f"https://api.github.com/repos/{repo}/contents/{path}?ref={branch}"
for attempt in range(3):
try:
with urllib.request.urlopen(urllib.request.Request(url, headers=HDRS), timeout=20) as r:
return ("ok", r.read().decode("utf-8", "replace"))
except urllib.error.HTTPError as e:
if e.code == 404:
break
if e.code in (403, 429):
time.sleep(60)
continue
time.sleep(2 * attempt)
except Exception:
time.sleep(2 * attempt)
return ("404", "")


def normal(d):
return (d or "").strip().rstrip("/").strip()


def valid_domain(d):
v = normal(d)
if not v or v in JUNK or v == DEFAULT_DOMAIN:
return False
if "ip6.arpa" in v or v == "free.hr" or v.endswith(".free.hr") or v == "cho.zzx.free":
return False
return True


def candidate_repos():
repos = {}
p1 = os.path.join(ROOT, "argo_all_results.json")
p2 = os.path.join(ROOT, "argo_indirect_results.json")
if os.path.exists(p1):
for r in json.load(open(p1, encoding="utf-8-sig")):
dom = r.get("files_idx_fallback") or r.get("root_idx_fallback")
if valid_domain(dom):
repos[r["repo"]] = normal(dom)
if os.path.exists(p2):
for r in json.load(open(p2, encoding="utf-8-sig")):
dom = r.get("files_index") or r.get("root_index")
if valid_domain(dom):
repos.setdefault(r["repo"], normal(dom))
return repos


def main():
repos = candidate_repos()
cache_path = os.path.join(ROOT, "file_content_cache.json")
cache = json.load(open(cache_path, encoding="utf-8-sig")) if os.path.exists(cache_path) else {}
print(f"== (11) UUID 提取: {len(repos)} 候选仓库(缓存命中 {sum(1 for r in repos if r in cache)})==")

out = {}

def from_cache(repo):
for path in ("files/index.js", "index.js"):
c = cache.get(repo, {}).get(path)
if c:
m = UUID_RE.search(c)
if m:
return m.group(1)
return None

def fetch_uuid(repo):
for path in ("files/index.js", "index.js"):
st, c = get_file(repo, path)
if st == "ok":
m = UUID_RE.search(c)
if m:
return m.group(1)
elif st != "404":
return f"error: {st}"
return "NOT_FOUND"

todo = []
for repo in repos:
u = from_cache(repo)
if u:
out[repo] = u
else:
todo.append(repo)

if todo:
print(f" 缓存未命中 {len(todo)},API 补抓 ...")
with concurrent.futures.ThreadPoolExecutor(max_workers=12) as ex:
futs = {ex.submit(fetch_uuid, r): r for r in todo}
for i, fut in enumerate(concurrent.futures.as_completed(futs), 1):
out[futs[fut]] = fut.result()
if i % 20 == 0:
print(f" {i}/{len(todo)}")

with open(os.path.join(ROOT, "uuid_results.json"), "w", encoding="utf-8") as f:
json.dump(out, f, ensure_ascii=False, indent=2)
ok = sum(1 for v in out.values() if re.fullmatch(r"[0-9a-fA-F-]{36}", v or ""))
print(f"完成:提取到 UUID {ok} / {len(out)}")


if __name__ == "__main__":
main()



14.13 ⑫ merge_final.py

# -*- coding: utf-8 -*-
"""⑫ 收尾合并(离线确定性操作,不发起任何网络请求)
1. 直接 400 清单 <- bad_request_results.json (kind == bad_request)
2. 间接 400 清单 <- probe_indirect_results.json (kind == 400_BAD_REQUEST)
3. 重建 bad_request_all.json {total, direct_count, indirect_count, overlap, domains}
4. 归并间接自定义域名到 custom_domains.txt(JUNK 过滤 + 剔除默认值 choreo.zzx.free.hr)
"""
import os, json, sys

try:
sys.stdout.reconfigure(encoding="utf-8", errors="replace")
except Exception:
pass

ROOT = os.path.dirname(os.path.abspath(__file__))
JUNK = {"", "NOT_FOUND", "NO_MATCH", "OBFUSCATED", "error"}
DEFAULT_DOMAIN = "choreo.zzx.free.hr"


def load_json(fn):
return json.load(open(os.path.join(ROOT, fn), encoding="utf-8-sig"))


def normal(domain):
return domain.strip().rstrip("/").strip()


def valid_indirect_domain(d):
v = normal(d or "")
if not v or v in JUNK or v == DEFAULT_DOMAIN:
return False
if "ip6.arpa" in v or v == "free.hr" or v.endswith(".free.hr") or v == "cho.zzx.free":
return False
return True


def main():
# 1) 直接 400
direct_400 = sorted({normal(i["domain"])
for i in load_json("bad_request_results.json").get("bad_request", [])
if normal(i["domain"])})
# 2) 间接 400
indirect_400 = sorted({normal(i["domain"])
for i in load_json("probe_indirect_results.json")
if i.get("kind") == "400_BAD_REQUEST"})
# 3) 合并
overlap = sorted(set(direct_400) & set(indirect_400))
merged = sorted(set(direct_400) | set(indirect_400))
out = {"total": len(merged),
"direct_count": len(direct_400),
"indirect_count": len(indirect_400),
"overlap": overlap,
"domains": merged}
with open(os.path.join(ROOT, "bad_request_all.json"), "w", encoding="utf-8") as f:
json.dump(out, f, ensure_ascii=False, indent=2)
print(f"bad_request_all.json: total={len(merged)} direct={len(direct_400)} "
f"indirect={len(indirect_400)} overlap={overlap}")

# 4) 归并间接自定义域名
cp = os.path.join(ROOT, "custom_domains.txt")
cur = {normal(l) for l in open(cp, encoding="utf-8-sig")} if os.path.exists(cp) else set()
cur.discard("")
ind = set()
for r in load_json("argo_indirect_results.json"):
for k in ("files_index", "root_index"):
v = r.get(k)
if valid_indirect_domain(v):
ind.add(normal(v))
new = sorted(cur | ind)
with open(cp, "w", encoding="utf-8") as f:
f.write("\n".join(new) + "\n")
print(f"custom_domains.txt: 原 {len(cur)} -> 归并后 {len(new)}(间接新增 {len(set(new) - cur)})")


if __name__ == "__main__":
main()



14.14 ⑬ build_relations.py

# -*- coding: utf-8 -*-
"""⑬ 构建 repo↔domain↔uuid 关系
输入: argo_all_results.json + argo_indirect_results.json(域名候选)、uuid_results.json(UUID)
输出:
repo_domain_uuid.json repo -> {domain, uuid}
repo_domains_map.json repo -> domain
rel_repo_uuid.json repo -> uuid
rel_domain_repo.json domain -> [repo, ...]
rel_domain_uuid.json domain -> uuid
"""
import os, re, json, sys

try:
sys.stdout.reconfigure(encoding="utf-8", errors="replace")
except Exception:
pass

ROOT = os.path.dirname(os.path.abspath(__file__))
JUNK = {"", "NOT_FOUND", "NO_MATCH", "OBFUSCATED", "error"}
BAD_DOMAINS = JUNK | {"choreo.zzx.free.hr", "free.hr", "cho.zzx.free"}
UUID_OK = re.compile(r"^[0-9a-fA-F]{8}-[0-9a-fA-F]{4}-[0-9a-fA-F]{4}-[0-9a-fA-F]{4}-[0-9a-fA-F]{12}$")


def load(fn):
return json.load(open(os.path.join(ROOT, fn), encoding="utf-8-sig"))


def normal(d):
return (d or "").strip().rstrip("/").strip()


def is_bad_domain(d):
v = normal(d)
if not v or v in BAD_DOMAINS:
return True
if "ip6.arpa" in v:
return True
if v == "free.hr" or v.endswith(".free.hr"):
return True
return False


def pick_domain(r):
for k in ("files_idx_fallback", "files_index", "root_idx_fallback", "root_index"):
v = normal(r.get(k))
if v and not is_bad_domain(v):
return v
return None


def dump(fn, obj):
with open(os.path.join(ROOT, fn), "w", encoding="utf-8") as f:
json.dump(obj, f, ensure_ascii=False, indent=2)
print(f" wrote {fn} ({len(obj)} 条)")


def main():
uuids = load("uuid_results.json")
records = {}
for src in ("argo_all_results.json", "argo_indirect_results.json"):
for r in load(src):
repo = r["repo"]
dom = pick_domain(r)
uu = uuids.get(repo)
if dom and uu and UUID_OK.match(uu):
records[repo] = {"domain": dom, "uuid": uu}

repo_domain_uuid = dict(sorted(records.items()))
repo_domains_map = {k: v["domain"] for k, v in repo_domain_uuid.items()}
rel_repo_uuid = {k: v["uuid"] for k, v in repo_domain_uuid.items()}

d2repos, d2uuid = {}, {}
for repo, v in repo_domain_uuid.items():
d2repos.setdefault(v["domain"], []).append(repo)
d2uuid[v["domain"]] = v["uuid"]
rel_domain_repo = {d: sorted(rs) for d, rs in sorted(d2repos.items())}
rel_domain_uuid = dict(sorted(d2uuid.items()))

dump("repo_domain_uuid.json", repo_domain_uuid)
dump("repo_domains_map.json", repo_domains_map)
dump("rel_repo_uuid.json", rel_repo_uuid)
dump("rel_domain_repo.json", rel_domain_repo)
dump("rel_domain_uuid.json", rel_domain_uuid)

shared = {d: rs for d, rs in rel_domain_repo.items() if len(rs) > 1}
if shared:
print("同一域名多 fork 共享:")
for d, rs in shared.items():
print(f" {d} <- {', '.join(rs)} -> {rel_domain_uuid[d]}")


if __name__ == "__main__":
main()



14.15 ⑭ gen_vless.py

# -*- coding: utf-8 -*-
# 生成 vless 链接:uuid / host / sni / 名称(Choreo-用户名) 按数据替换(全量,不跳过 400 域名)
import json

rdu = json.load(open('repo_domain_uuid.json', encoding='utf-8-sig'))

tpl = ('vless://{uuid}@www.visa.com.tw:443?encryption=none&security=tls'
'&sni={domain}&type=ws&host={domain}&path=%2F#Choreo-{user}')

links = []
for repo, info in sorted(rdu.items()):
domain, uuid = info.get('domain'), info.get('uuid')
if not domain or not uuid:
continue
user = repo.split('/')[0]
links.append(tpl.format(uuid=uuid, domain=domain, user=user))

seen, uniq = set(), []
for l in links:
if l not in seen:
seen.add(l); uniq.append(l)

open('vless_links.txt', 'w', encoding='utf-8').write('\n'.join(uniq) + '\n')
print(f'生成 {len(uniq)} 条链接(去重前 {len(links)})-> vless_links.txt')
print('示例:', uniq[0] if uniq else '无')




15. VLESS 链接生成说明(⑭)

模板(用户提供):

vless://[email protected]:443?encryption=none&security=tls&sni=choreo-2.ncvfjhgr.gq&type=ws&host=choreo-2.ncvfjhgr.gq&path=%2F#Choreo-ardjsrset



替换规则(gen_vless.py 按 repo_domain_uuid.json 全量生成,不跳过 400 域名):

模板位置 替换为
525476d1-...-d53c30935d71 该 repo 提取到的 uuid
choreo-2.ncvfjhgr.gq(host 与 sni 两处) 该 repo 提取到的 domain
ardjsrset(名称后缀) 仓库 owner(用户名),不含仓库名
地址 www.visa.com.tw:443、type=ws、path=%2F、encryption=none、security=tls 固定保留

输出:vless_links.txt,351 条(按整条链接去重;同一 UUID 被多 fork 共享时域名不同仍各占一条)。

核验样例:

vless://[email protected]:443?encryption=none&security=tls&sni=choreo0402.moisturize45.com&type=ws&host=choreo0402.moisturize45.com&path=%2F#Choreo-DLLJ34
vless://[email protected]:443?encryption=none&security=tls&sni=choreo.wto.pp.ua&type=ws&host=choreo.wto.pp.ua&path=%2F#Choreo-xuanyepan




14. 附录:完整脚本归档(2026-09-21 重建验证版,可直接复制运行)

以下 15 个脚本即最新一次全量重建实际使用的版本(582 直接 fork → 796 全量 → 332 域名 → 351 UUID → 351 条 VLESS)。
复现方式:在空目录按 §13.5 顺序保存并执行即可。所有脚本自动读 GITHUB_TOKEN 环境变量;
统一 utf-8-sig 读取历史 JSON、sys.stdout.reconfigure 规避 PowerShell 中文坑;
⑥⑦⑪ 走 api.github.com contents API(绕开 raw CDN 卡顿)。

14.1 ① fetch_direct_forks.py

# -*- coding: utf-8 -*-
"""① 枚举 eooce/choreo-2go 直接 fork(GitHub REST API 分页,单页失败重试 5 次带退避)
输出: forks_direct_full.json / forks.json / forks_all.json / direct_forks.txt / all_forks.txt
"""
import os, json, time, sys, urllib.request, urllib.error

try:
sys.stdout.reconfigure(encoding="utf-8", errors="replace")
except Exception:
pass

ROOT = os.path.dirname(os.path.abspath(__file__))
OWNER_REPO = "eooce/choreo-2go"
TOKEN = os.environ.get("GITHUB_TOKEN", "")
HDRS = {"User-Agent": "Mozilla/5.0 (CodeArtsWork collect)",
"Accept": "application/vnd.github+json"}
if TOKEN:
HDRS["Authorization"] = f"token {TOKEN}"


def api_get(url):
req = urllib.request.Request(url, headers=HDRS)
with urllib.request.urlopen(req, timeout=30) as r:
rem = r.headers.get("X-RateLimit-Remaining")
if rem:
print(f" [rate] remaining={rem}/{r.headers.get('X-RateLimit-Limit')}")
return r.read().decode("utf-8", "replace")


def fetch_all_forks():
forks, page = [], 1
while True:
url = f"https://api.github.com/repos/{OWNER_REPO}/forks?per_page=100&page={page}"
data = None
for attempt in range(1, 6):
try:
data = json.loads(api_get(url))
break
except urllib.error.HTTPError as e:
if e.code == 403:
reset = e.headers.get("X-RateLimit-Reset")
wait = max(30, int(reset) - int(time.time()) + 5) if reset else 120
print(f" [403 rate-limit] sleep {min(wait, 600)}s (attempt {attempt}/5)")
time.sleep(min(wait, 600))
else:
print(f" [http {e.code}] retry {attempt}/5 ...")
time.sleep(5 * attempt)
except Exception as e:
print(f" [err] {e!r} retry {attempt}/5 ...")
time.sleep(5 * attempt)
if not data:
raise RuntimeError(f"page {page} failed after retries")
forks.extend(data)
print(f" page {page}: +{len(data)} (total {len(forks)})")
if len(data) < 100:
break
page += 1
return forks


def trim(f):
return {"full_name": f.get("full_name"),
"fork": f.get("fork"),
"forks_count": f.get("forks_count"),
"created_at": f.get("created_at"),
"pushed_at": f.get("pushed_at"),
"default_branch": f.get("default_branch")}


def main():
print(f"== (1) 枚举直接 fork: {OWNER_REPO} ==")
forks = fetch_all_forks()
trimmed = [trim(f) for f in forks]
names = sorted({t["full_name"] for t in trimmed if t["full_name"]})

def dump(fn, obj):
with open(os.path.join(ROOT, fn), "w", encoding="utf-8") as f:
json.dump(obj, f, ensure_ascii=False, indent=2)
print(f" wrote {fn}")

dump("forks_direct_full.json", trimmed)
dump("forks.json", trimmed)
dump("forks_all.json", trimmed)
with open(os.path.join(ROOT, "direct_forks.txt"), "w", encoding="utf-8") as f:
f.write("\n".join(names) + "\n")
with open(os.path.join(ROOT, "all_forks.txt"), "w", encoding="utf-8") as f:
f.write("\n".join(names) + "\n")
total_sub = sum(t["forks_count"] or 0 for t in trimmed)
print(f"直接 fork 总数: {len(names)};子 fork 数合计: {total_sub}")


if __name__ == "__main__":
main()



14.2 ② fetch_network_members.py

# -*- coding: utf-8 -*-
"""② 抓取 network/members 全部分页 HTML,合并保存为 net_dump.html,
并提取 repo 链接(network_members.txt)与 hovercard owner(network_members_owners.txt)"""
import os, re, sys, time, urllib.request

try:
sys.stdout.reconfigure(encoding="utf-8", errors="replace")
except Exception:
pass

ROOT = os.path.dirname(os.path.abspath(__file__))
ROOT_REPO = "eooce/choreo-2go"
TOKEN = os.environ.get("GITHUB_TOKEN", "")
HDRS = {"User-Agent": "Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 "
"(KHTML, like Gecko) Chrome/124.0 Safari/537.36",
"Accept": "text/html,application/xhtml+xml"}
if TOKEN:
HDRS["Authorization"] = f"token {TOKEN}"

RE_REPO = re.compile(r'href="/([^/"]+/[^/"]+)"')
RE_HOVER = re.compile(r'data-hovercard-url="/users/([^/"]+)/hovercard"')


def fetch(url):
for a in range(1, 6):
try:
req = urllib.request.Request(url, headers=HDRS)
with urllib.request.urlopen(req, timeout=30) as r:
return r.read().decode("utf-8", "replace")
except Exception as e:
print(f" [retry {a}/5] {e!r}")
time.sleep(3 * a)
raise RuntimeError("fetch failed: " + url)


def main():
print(f"== (2) network members: {ROOT_REPO} ==")
all_repos, all_owners = set(), set()
page, stagnant = 1, 0
dump_parts = []
while page <= 60 and stagnant < 2:
base = f"https://github.com/{ROOT_REPO}/network/members"
url = base if page == 1 else f"{base}?page={page}"
html = fetch(url)
dump_parts.append(f"<!-- ===== page {page} ===== -->\n{html}")
repos = {m.group(1).strip("/") for m in RE_REPO.finditer(html)}
owners = set(RE_HOVER.findall(html))
new = repos - all_repos
print(f" page {page}: repos={len(repos)} new={len(new)} owners_total={len(all_owners | owners)}")
all_repos |= repos
all_owners |= owners
stagnant = stagnant + 1 if not new else 0
page += 1
time.sleep(1)

with open(os.path.join(ROOT, "net_dump.html"), "w", encoding="utf-8") as f:
f.write("\n".join(dump_parts))
with open(os.path.join(ROOT, "network_members.txt"), "w", encoding="utf-8") as f:
f.write("\n".join(sorted(all_repos)) + "\n")
with open(os.path.join(ROOT, "network_members_owners.txt"), "w", encoding="utf-8") as f:
f.write("\n".join(sorted(all_owners)) + "\n")
print(f"network members 粗提取: {len(all_repos)} repos / {len(all_owners)} owners(含干扰,稍后由 extract_network 清洗)")


if __name__ == "__main__":
main()



14.3 ③ analyze_net.py

# -*- coding: utf-8 -*-
"""③ 分析 net_dump.html 结构(可选步骤)"""
import os, re, sys

try:
sys.stdout.reconfigure(encoding="utf-8", errors="replace")
except Exception:
pass

ROOT = os.path.dirname(os.path.abspath(__file__))
html = open(os.path.join(ROOT, "net_dump.html"), encoding="utf-8", errors="replace").read()

print("size:", len(html))
print("repo hrefs:", len(re.findall(r'href="/[^/"]+/[^/"]+"', html)))
print("hovercards:", len(re.findall(r'data-hovercard-url="/users/', html)))

pages = sorted(set(re.findall(r'href="([^"]*network/members[^"]*)"', html)))
print("page links:", pages if len(pages) < 20 else f"{len(pages)} links (sample {pages[:5]})")

hover = re.findall(r'data-hovercard-url="([^"]+)"', html)
print("hovercard sample:", hover[:5])

m = re.search(r'aria-label="Pagination".{0,4000}?</nav>', html, re.S)
print("pagination block:", "found" if m else "not found")



14.4 ④ extract_network.py

# -*- coding: utf-8 -*-
"""④ 从 net_dump.html 提取 repo key
输出: net_repo_keys.json / indirect_forks.txt / indirect_forks_clean.txt
(清洗:排除 root 仓库、非 owner/repo 形态、导航/静态资源干扰链接)"""
import os, re, json, sys

try:
sys.stdout.reconfigure(encoding="utf-8", errors="replace")
except Exception:
pass

ROOT = os.path.dirname(os.path.abspath(__file__))
ROOT_REPO = "eooce/choreo-2go"

html = open(os.path.join(ROOT, "net_dump.html"), encoding="utf-8", errors="replace").read()
raw = set(re.findall(r'href="/([^/"]+/[^/"]+)"', html))

BANNED_PREFIX = ("features/", "settings/", "topics/", "collections/", "notifications",
"login", "join", "explore", "apps/", "orgs/", "account/", "marketplace",
"pulls", "issues", "search", "security", "pricing", "customer-stories",
"readme", "about", "site/", "enterprise", "team", "sponsors", "new",
"codespaces", "generative-ai", "contact", "terms", "privacy")
STATIC_EXT = (".js", ".json", ".md", ".py", ".html", ".css", ".png", ".svg", ".txt")


def ok(r):
if not re.fullmatch(r"[\w.\-]+/[\w.\-]+", r):
return False
if r.lower().endswith(STATIC_EXT):
return False
if any(r.startswith(b) or r.lower().startswith(b) for b in BANNED_PREFIX):
return False
return True


keys = sorted({r for r in raw if ok(r) and r.lower() != ROOT_REPO.lower()})

with open(os.path.join(ROOT, "net_repo_keys.json"), "w", encoding="utf-8") as f:
json.dump(keys, f, ensure_ascii=False, indent=2)
with open(os.path.join(ROOT, "indirect_forks.txt"), "w", encoding="utf-8") as f:
f.write("\n".join(keys) + "\n")
with open(os.path.join(ROOT, "indirect_forks_clean.txt"), "w", encoding="utf-8") as f:
f.write("\n".join(keys) + "\n")

print(f"原始 href 候选: {len(raw)};清洗后间接 fork: {len(keys)}")



14.5 ⑤ check_argo.py

# -*- coding: utf-8 -*-
"""⑤ 初轮 ARGO 域名提取(93 仓库)
原始版本硬编码 93 个 owner/repo;重建版改为取 direct_forks.txt 排序后前 93 个(确定性等价)。
输出: argo_results.json
"""
import os, re, json, sys, time, urllib.request, urllib.error, concurrent.futures

try:
sys.stdout.reconfigure(encoding="utf-8", errors="replace")
except Exception:
pass

ROOT = os.path.dirname(os.path.abspath(__file__))
ARGO_RE = re.compile(r"ARGO_DOMAIN\s*=\s*process\.env\.ARGO_DOMAIN\s*\|\|\s*(['\"])(.*?)\1")
HDRS = {"User-Agent": "Mozilla/5.0 (CodeArtsWork check)",
"Accept": "application/vnd.github.raw"}
if os.environ.get("GITHUB_TOKEN"):
HDRS["Authorization"] = f"token {os.environ['GITHUB_TOKEN']}"


def get_file(repo, path):
"""api.github.com contents API(raw CDN 在本环境卡顿),main→master 回退"""
for branch in ("main", "master"):
url = f"https://api.github.com/repos/{repo}/contents/{path}?ref={branch}"
for attempt in range(3):
try:
with urllib.request.urlopen(urllib.request.Request(url, headers=HDRS), timeout=20) as r:
return ("ok", r.read().decode("utf-8", "replace"))
except urllib.error.HTTPError as e:
if e.code == 404:
break
if e.code in (403, 429):
time.sleep(60)
continue
time.sleep(2 * attempt)
except Exception:
time.sleep(2 * attempt)
return ("404", "")


def check_repo(repo):
res = {"repo": repo}
st1, c1 = get_file(repo, "files/index.js") # 首选 files/index.js
m1 = ARGO_RE.search(c1) if st1 == "ok" else None
res["files_idx_fallback"] = m1.group(2) if m1 else None
st2, c2 = get_file(repo, "index.js") # 次选根 index.js
m2 = ARGO_RE.search(c2) if st2 == "ok" else None
res["root_idx_fallback"] = m2.group(2) if m2 else None
return res


def main():
repos = [l.strip() for l in open(os.path.join(ROOT, "direct_forks.txt"),
encoding="utf-8-sig") if l.strip()]
repos = sorted(repos)[:93]
print(f"== (5) 初轮 ARGO 检测: {len(repos)} 仓库 ==")
out = []
with concurrent.futures.ThreadPoolExecutor(max_workers=12) as ex:
futs = {ex.submit(check_repo, r): r for r in repos}
for i, fut in enumerate(concurrent.futures.as_completed(futs), 1):
out.append(fut.result())
if i % 20 == 0:
print(f" {i}/{len(repos)}")
out.sort(key=lambda r: r["repo"])
with open(os.path.join(ROOT, "argo_results.json"), "w", encoding="utf-8") as f:
json.dump(out, f, ensure_ascii=False, indent=2)
ff = [r for r in out if r["files_idx_fallback"]]
rf = [r for r in out if r["root_idx_fallback"]]
nf = [r for r in out if not r["files_idx_fallback"] and not r["root_idx_fallback"]]
print(f"files/index.js 非空 fallback: {len(ff)};根 index.js 非空: {len(rf)};空/缺失: {len(nf)}")


if __name__ == "__main__":
main()



14.6 ⑥ check_all_forks.py

# -*- coding: utf-8 -*-
"""⑥ 全部 fork 的 ARGO 域名提取 → argo_all_results.json
附带:汇总直接 fork 的自定义域名 → custom_domains.txt(junk 过滤 + 剔除默认域名)。
文件抓取走 api.github.com contents API(raw.githubusercontent.com 在本环境严重卡顿);
抓到的 index.js 内容写入 file_content_cache.json 供 ⑪ fetch_uuid 复用(省 API 配额)。
注意:所有读取方统一用 encoding="utf-8-sig"(历史版本首条记录带 BOM)。
"""
import os, re, json, sys, time, urllib.request, urllib.error, concurrent.futures

try:
sys.stdout.reconfigure(encoding="utf-8", errors="replace")
except Exception:
pass

ROOT = os.path.dirname(os.path.abspath(__file__))
ARGO_RE = re.compile(r"ARGO_DOMAIN\s*=\s*process\.env\.ARGO_DOMAIN\s*\|\|\s*(['\"])(.*?)\1")
JUNK = {"", "NOT_FOUND", "NO_MATCH", "OBFUSCATED", "error"}
DEFAULT_DOMAIN = "choreo.zzx.free.hr"
TOKEN = os.environ.get("GITHUB_TOKEN", "")
HDRS = {"User-Agent": "Mozilla/5.0 (CodeArtsWork check)",
"Accept": "application/vnd.github.raw"}
if TOKEN:
HDRS["Authorization"] = f"token {TOKEN}"


class RateLimited(Exception):
pass


def get_file(repo, path):
"""api.github.com contents API,main→master 回退,重试 3 次"""
for branch in ("main", "master"):
url = f"https://api.github.com/repos/{repo}/contents/{path}?ref={branch}"
for attempt in range(3):
try:
with urllib.request.urlopen(urllib.request.Request(url, headers=HDRS), timeout=20) as r:
rem = r.headers.get("X-RateLimit-Remaining")
if rem == "0":
raise RateLimited("remaining=0")
return ("ok", r.read().decode("utf-8", "replace"))
except urllib.error.HTTPError as e:
if e.code == 404:
break # 下一 branch
if e.code in (403, 429):
reset = e.headers.get("X-RateLimit-Reset")
wait = min(600, max(20, int(reset) - int(time.time()) + 5)) if reset else 120
print(f" [rate-limit] sleep {wait}s")
time.sleep(wait)
continue
time.sleep(2 * attempt)
except RateLimited:
reset = None
try:
with urllib.request.urlopen(urllib.request.Request(
"https://api.github.com/rate_limit", headers=HDRS), timeout=15) as r:
reset = json.loads(r.read())["resources"]["core"]["reset"]
except Exception:
pass
wait = min(600, max(20, int(reset) - int(time.time()) + 5)) if reset else 120
print(f" [rate-limit remaining=0] sleep {wait}s")
time.sleep(wait)
except Exception:
time.sleep(2 * attempt)
return ("404", "")


def check_repo(repo):
res = {"repo": repo}
st1, c1 = get_file(repo, "files/index.js")
m1 = ARGO_RE.search(c1) if st1 == "ok" else None
res["files_idx_fallback"] = m1.group(2) if m1 else None
st2, c2 = get_file(repo, "index.js")
m2 = ARGO_RE.search(c2) if st2 == "ok" else None
res["root_idx_fallback"] = m2.group(2) if m2 else None
res["_cache"] = {}
if st1 == "ok":
res["_cache"]["files/index.js"] = c1
if st2 == "ok":
res["_cache"]["index.js"] = c2
return res


def valid_domain(d):
if not d:
return False
d = d.strip().rstrip("/").strip()
if d in JUNK or d == DEFAULT_DOMAIN:
return False
if "ip6.arpa" in d or d == "free.hr" or d.endswith(".free.hr") or d == "cho.zzx.free":
return False
return True


def main():
data = json.load(open(os.path.join(ROOT, "forks_all.json"), encoding="utf-8-sig"))
repos = sorted({d["full_name"] for d in data if d.get("full_name")})
print(f"== (6) 全部 fork ARGO 检测: {len(repos)} 仓库 ==")
out = []
with concurrent.futures.ThreadPoolExecutor(max_workers=12) as ex:
futs = {ex.submit(check_repo, r): r for r in repos}
for i, fut in enumerate(concurrent.futures.as_completed(futs), 1):
out.append(fut.result())
if i % 50 == 0:
print(f" {i}/{len(repos)}")
out.sort(key=lambda r: r["repo"])
# 拆缓存:argo_all_results 只留元数据,文件内容进 file_content_cache.json
cache = {}
for r in out:
c = r.pop("_cache", {})
if c:
cache[r["repo"]] = c
with open(os.path.join(ROOT, "argo_all_results.json"), "w", encoding="utf-8") as f:
json.dump(out, f, ensure_ascii=False, indent=2)
with open(os.path.join(ROOT, "file_content_cache.json"), "w", encoding="utf-8") as f:
json.dump(cache, f, ensure_ascii=False)

doms = set()
for r in out:
for k in ("files_idx_fallback", "root_idx_fallback"):
v = r.get(k)
if valid_domain(v):
doms.add(v.strip().rstrip("/").strip())
with open(os.path.join(ROOT, "custom_domains.txt"), "w", encoding="utf-8") as f:
f.write("\n".join(sorted(doms)) + "\n")

ff = sum(1 for r in out if r["files_idx_fallback"])
rf = sum(1 for r in out if r["root_idx_fallback"])
print(f"完成:files 非空 {ff} / root 非空 {rf} / custom_domains.txt {len(doms)} 域名")


if __name__ == "__main__":
main()



14.7 ⑦ check_indirect_forks.py

# -*- coding: utf-8 -*-
"""⑦ 间接 fork(network members)ARGO 检测 → argo_indirect_results.json
每条记录 files_index / root_index 两个 fallback 候选。
文件抓取走 api.github.com contents API;内容并入 file_content_cache.json 供 ⑪ 复用。
输入: indirect_forks_clean.txt(extract_network.py 产出),并与 net_repo_keys.json 取并集。
"""
import os, re, json, sys, time, urllib.request, urllib.error, concurrent.futures

try:
sys.stdout.reconfigure(encoding="utf-8", errors="replace")
except Exception:
pass

ROOT = os.path.dirname(os.path.abspath(__file__))
ARGO_RE = re.compile(r"ARGO_DOMAIN\s*=\s*process\.env\.ARGO_DOMAIN\s*\|\|\s*(['\"])(.*?)\1")
ROOT_REPO = "eooce/choreo-2go"
TOKEN = os.environ.get("GITHUB_TOKEN", "")
HDRS = {"User-Agent": "Mozilla/5.0 (CodeArtsWork check)",
"Accept": "application/vnd.github.raw"}
if TOKEN:
HDRS["Authorization"] = f"token {TOKEN}"


def get_file(repo, path):
"""api.github.com contents API,main→master 回退,重试 3 次"""
for branch in ("main", "master"):
url = f"https://api.github.com/repos/{repo}/contents/{path}?ref={branch}"
for attempt in range(3):
try:
with urllib.request.urlopen(urllib.request.Request(url, headers=HDRS), timeout=20) as r:
rem = r.headers.get("X-RateLimit-Remaining")
if rem == "0":
raise RateLimited("remaining=0")
return ("ok", r.read().decode("utf-8", "replace"))
except urllib.error.HTTPError as e:
if e.code == 404:
break
if e.code in (403, 429):
reset = e.headers.get("X-RateLimit-Reset")
wait = min(600, max(20, int(reset) - int(time.time()) + 5)) if reset else 120
print(f" [rate-limit] sleep {wait}s")
time.sleep(wait)
continue
time.sleep(2 * attempt)
except RateLimited:
reset = None
try:
with urllib.request.urlopen(urllib.request.Request(
"https://api.github.com/rate_limit", headers=HDRS), timeout=15) as r:
reset = json.loads(r.read())["resources"]["core"]["reset"]
except Exception:
pass
wait = min(600, max(20, int(reset) - int(time.time()) + 5)) if reset else 120
print(f" [rate-limit remaining=0] sleep {wait}s")
time.sleep(wait)
except Exception:
time.sleep(2 * attempt)
return ("404", "")


class RateLimited(Exception):
pass


def check_repo(repo):
res = {"repo": repo}
st1, c1 = get_file(repo, "files/index.js")
m1 = ARGO_RE.search(c1) if st1 == "ok" else None
res["files_index"] = m1.group(2) if m1 else None
st2, c2 = get_file(repo, "index.js")
m2 = ARGO_RE.search(c2) if st2 == "ok" else None
res["root_index"] = m2.group(2) if m2 else None
res["_cache"] = {}
if st1 == "ok":
res["_cache"]["files/index.js"] = c1
if st2 == "ok":
res["_cache"]["index.js"] = c2
return res


def main():
repos = set()
p1 = os.path.join(ROOT, "indirect_forks_clean.txt")
p2 = os.path.join(ROOT, "net_repo_keys.json")
if os.path.exists(p1):
repos |= {l.strip() for l in open(p1, encoding="utf-8-sig") if l.strip()}
if os.path.exists(p2):
repos |= set(json.load(open(p2, encoding="utf-8-sig")))
repos = sorted(r for r in repos if r and r.lower() != ROOT_REPO.lower())
print(f"== (7) 间接 fork ARGO 检测: {len(repos)} 仓库 ==")
out = []
with concurrent.futures.ThreadPoolExecutor(max_workers=12) as ex:
futs = {ex.submit(check_repo, r): r for r in repos}
for i, fut in enumerate(concurrent.futures.as_completed(futs), 1):
out.append(fut.result())
if i % 50 == 0:
print(f" {i}/{len(repos)}")
out.sort(key=lambda r: r["repo"])
cache_path = os.path.join(ROOT, "file_content_cache.json")
cache = json.load(open(cache_path, encoding="utf-8-sig")) if os.path.exists(cache_path) else {}
for r in out:
c = r.pop("_cache", {})
if c:
cache.setdefault(r["repo"], {}).update(c)
with open(os.path.join(ROOT, "argo_indirect_results.json"), "w", encoding="utf-8") as f:
json.dump(out, f, ensure_ascii=False, indent=2)
with open(cache_path, "w", encoding="utf-8") as f:
json.dump(cache, f, ensure_ascii=False)
fi = sum(1 for r in out if r["files_index"])
ri = sum(1 for r in out if r["root_index"])
print(f"完成:files_index 非空 {fi} / root_index 非空 {ri}")


if __name__ == "__main__":
main()