同时运行多个网关

在一台机器上,把多个 profile——各自带自己的 bot token、会话和记忆——作为托管服务运行。本页讲运维事项:一起启动它们、跨 profile 看日志、防止主机休眠,以及从常见 launchd/systemd 怪癖中恢复。

如果你只跑一个 Hermes agent,不需要本页——基础见 Profiles。如果你的实例在不同机器上、一个桌面应用要同时访问它们,见把桌面端连接到多个 Hermes 实例。

何时用这个

当你有两个或更多 Hermes agent 应同时在线时,你需要这套设置。常见原因:

  • 一个 Telegram bot 上跑个人助理,另一个上跑编码 agent
  • 每个家庭成员一个 agent,或每个 Slack 工作区一个
  • 同一配置的沙箱 + 生产实例
  • 研究 agent + 写作 agent + cron 驱动的 bot——各自隔离记忆和 skill

每个 profile 已经有自己的按平台 supervisor 条目:一个 LaunchAgent(ai.hermes.gateway-<name>.plist)、一个 systemd 用户服务(hermes-gateway-<name>.service)、用 sudo hermes gateway install --system 安装时的 systemd 系统服务(经 User= 以调用用户身份运行)、Windows 计划任务,或 s6/Docker 服务——桌面应用还会派生自己的按 profile hermes serve 后端。本指南补充集体管理它们的模式。

快速上手

# 创建 profiles(一次)
hermes profile create coder
hermes profile create personal-bot
hermes profile create research

# 配置每个
coder setup
personal-bot setup
research setup

# 把每个网关装为托管服务
coder gateway install
personal-bot gateway install
research gateway install

# 启动它们全部
coder gateway start
personal-bot gateway start
research gateway start

就这样——三个独立 agent,各跑在自己的进程里,崩溃时和用户登录时自动重启。

另一种:一个网关服务所有 profile(多路复用)

上面的模型每个 profile 一个进程。另一种是单个多路复用网关:一个网关进程——无论哪个 profile 启动它——成为唯一的入站进程,为机器上每个 profile 服务消息。

默认 profile 的生命周期命令作用于那个进程。命名 profile 可以只停或重启自己的 bot,而不停宿主机:

  • 它在线时,hermes -p <name> gateway run attach 而非启动第二个进程:它打印宿主机网关的 PID 和服务集合,退出 0。如果 <name> 还没被服务,它请宿主机网关重扫 profiles/,一旦答复包含它就 attach;只有在无法让宿主机网关服务它时才拒绝(非零退出)。
  • hermes gateway start --all / restart --all 指的是那个宿主机多路复用器。它们绝不会清扫机器上每个网关进程;仍自跑网关的 profile 会被报告,绝不被杀,并附 hermes gateway migrate --multiplex 一行命令。
  • hermes gateway run --replace 接管服务这个 profile 的进程,无论它是哪个 profile 启动的。当宿主机属主是另一个 profile 的独立网关(未迁移的按 profile 机群)时,它从不服务这个 profile,因此 --replace 像普通 run 一样在它旁边启动,而不是拒绝并在 supervisor 下引发重生风暴。旧版 Hermes 写过一个 systemd drop-in(hermes-gateway.service.d/20-replace.conf)强制把 --replace 加到 unit;hermes update / hermes gateway restart 现在会移除那个文件。hermes gateway run --force 完全不询问宿主机进程就启动一个独立网关(宿主机卡住或应答错误时的逃生舱)。
  • 在服务 supervisor 下,attach 退出码是 75 而非 0——systemd、s6 和 launchd 都会在短暂延迟后重启 75,因此 unit 不断重试,并在宿主机进程消失时自己接管。
  • 两个同时启动的 unit 都可能还没看到宿主机进程;宿主机锁决定谁运行,败者退出 75 并在重试时 attach。--replace 不跳过该检查(每个生成的 unit 都带),只有 --force 跳过。

多路复用默认开启(gateway.multiplex_profiles 默认 true),带一条安全规则:未设置的标志是一个请求,由默认网关在启动时裁定,绝不是定论。每次启动它都运行与 [hermes gateway migrate --multiplex](#从按 profile 网关迁移) 相同的预检,只有折叠安全时才多路复用——两个或更多 profile、没有次要 profile 仍自跑网关(活跃进程或已装服务,或 s6 下真正起来的按 profile 槽位)、没有重复 bot 凭据、没有不带 /p/<profile>/ 入口的端口绑定平台。未设置即开启:没有阻塞时,网关多路复用并把 gateway.multiplex_profiles: true 写进默认 profile 的 config.yaml(保留注释),让文件反映运行时行为。否则它只服务默认 profile 并在有其他 profile 的宿主机上大声说明:网关启动时一个方框警告,点名未被服务的 profile、阻塞原因和修复;hermes update 摘要和 hermes gateway status 中同样的方框;仪表盘一条横幅(/api/status 带 multiplex_standalone_reason)。单 profile 安装不警告——没东西可服务。拒绝时什么都不写。

显式的 true 绕过迁移预检,但一个用 gateway.standalone: true 退出的启动 profile 除外:

  • gateway.multiplex_profiles: true(迁移写入的)无视预检多路复用——是你或迁移做的决定。
  • gateway.multiplex_profiles 现在只有一个合法值:true,并替你写好。未设置键解析为开启并在默认 profile 的 config.yaml 中显式化。false 已退役:网关就地把它改写成 true,并在那次启动和下次 hermes update 摘要中打印一次性方框通知——绝不静默翻转。一个按 profile 网关在该 profile 自己的配置里是 gateway.standalone: true(临时垫片,非受支持拓扑),或下面边界情况用 --force。
  • 进程环境中的 GATEWAY_MULTIPLEX_PROFILES 与显式 true 一样覆盖未设置键的裁定。
  • 命名 profile 自己 config.yaml(profiles/<name>/config.yaml)里的 gateway.standalone: true 是临时兼容垫片(见下文):宿主机网关不服务该 profile,该 profile 不用 --force 跑自己的网关。它的网关只服务自己,即使也设了 multiplex_profiles: true(见不再新建按 profile 网关)。设在默认 profile 上会被忽略并警告——默认 profile 就是宿主机网关。这个键没有环境变量。

其他进程(hermes -p <name> gateway start、仪表盘、hermes gateway migrate)绝不猜未设置标志如何裁定:它们读运行中默认网关的 served_profiles 记录,只在没有网关运行时才回退到显式标志。

何时优先多路复用

  • 容器/VPS 部署,N 个 supervisor unit、N 个端口、N 个 PID 文件是负担。
  • 许多低流量 profile,每个都不值得一个完整进程。
  • 你想要一个东西来启动、监控和重启。

每 profile 一进程不再是一个隐式选择的拓扑:命名 profile 的 gateway install / gateway start 不带 --force 会拒绝(见不再新建按 profile 网关)。多路复用缺口仍开着、却仍需要自己网关的 profile,可在自己 config.yaml 设临时 gateway.standalone: true;真实边界挡住折叠的地方——机群跨 UNIX 用户拆分,或 HERMES_HOME 在 <default home>/profiles/ 之外——每个 profile 以 --force 为路径。

固定标志

标志未设置时,默认网关每次启动裁定(见上)。要固定它,在网关作为宿主机进程运行的那个 profile(通常是默认 profile)上设置并重启其网关——true 强制多路复用,即使启动预检本会拦住(false 已退役并忽略):

hermes config set gateway.multiplex_profiles true
hermes gateway restart

等价地,在默认 profile 的 ~/.hermes/config.yaml:

gateway:
  multiplex_profiles: true

(为方便,也接受顶层 multiplex_profiles: true。)多路复用时,默认网关枚举每个 profile,用该 profile 自己的凭据拉起每个 profile 启用的平台,并把每条入站消息路由到它所属的 profile。每一轮都解析被路由 profile 的配置、skill、记忆、SOUL 和提供商 key——凭据绝不跨 profile 共享。

宿主机自动服务未停驻的次要 profile。在停驻的 profile 上用 gateway start 把它带回在线。

不停宿主机地停一个 profile

对于由宿主机多路复用器服务的命名 profile:

hermes -p coder gateway stop     # 停驻 coder;其他 profile 继续跑
hermes -p coder gateway start    # 取消停驻 coder 并再次服务它
hermes -p coder gateway restart  # 用当前配置重连 coder

stop 在请宿主机停止该 profile 的适配器并把其 cron 任务从后续 tick 中排除之前,先在 profile home 写 gateway.parked。标记跨宿主机重启持久。其内容被忽略;空文件即可。开通时可预创建 <profiles-root>/coder/gateway.parked,让一个已装 profile 保持离线。停驻不删除 profile、其会话或其计划任务。

start 移除标记,然后请运行中的宿主机服务该 profile。没有运行中的宿主机时它移除标记并走正常启动路径;如提示,从默认 profile 启动宿主机。restart 取消服务再服务该 profile,不写停驻标记,重读其配置。在一个停驻且无活跃按 profile 网关的 profile 上,restart 行为同 start:移除标记并热服务该 profile(标记旁用 --force 启动的网关保持自己的重启)。这些操作不终止 cron tick 已派发的工作。

宿主机还每 30 秒重扫:手动加标记会取消服务该 profile;手动移除它使其重新合格。如果控制 socket 不确认请求,CLI 会说明,下次重扫应用标记状态。适配器拆除或连接可能额外耗时。标记存在时 hermes -p coder gateway status 报 parked (hermes -p coder gateway start)。

启动 profile 不能被取消服务。默认 profile 的标记被忽略并警告;其生命周期命令和 --all 变体保持整机行为。单独运行的 --force 网关保持自己的进程生命周期。

仪表盘和桌面应用对已服务 profile 的 Stop / Start 按钮做同样的事:Stop 停驻它(/api/gateway/stop?profile=coder 派生 hermes -p coder gateway stop),Start 在宿主机网关在线时取消停驻,/api/status 列出 parked_profiles。对一个未停驻的命名 profile 点 Start 仍回 409——它需要自己的网关。

停驻 vs gateway.standalone: true

二者从不同时作用于同一 profile,且互不蕴含:

Profile宿主机服务它-p X gateway stop-p X gateway start
已服务(默认情形)是停驻它(标记 + unserve-profile)已服务
停驻(gateway.parked 存在)否,直到取消停驻已停驻取消停驻(移除标记 + serve-profile)
独立(gateway.standalone: true)从不停它自己的网关进程;不写标记启动它自己的网关

gateway.standalone 胜出:宿主机绝不只为独立 profile 写 gateway.parked,也从不服务它,无论停驻与否,因此该 profile 的 stop 和 start 保持按进程含义。停驻是多路复用原生的下线一个 profile 的方式——它补上了临时垫片一直留着的"按 profile 停/重启"缺口。

不再新建按 profile 网关

默认情况下,一个宿主机网关服务每个 profile,命名 profile 不获得自己的网关。不带下面的退出选项,hermes -p coder gateway install(或 start、run,以及 hermes -p coder setup 的服务步骤)无论当前是否有宿主机网关在跑都以退出码 78 拒绝:

❌ Profile 'coder' 不获得自己的网关。
  每宿主机恰好一个网关进程是所有 profile 的入站进程。
  为这个 profile 启动独立网关会双重绑定其平台
  (一个 bot token 上两个轮询器、端口冲突)。

  从默认 profile 安装或启动宿主机网关;它也服务这一个:

    hermes gateway install

  或把已有按 profile 机群折叠到一个宿主机网关:

    hermes gateway migrate --multiplex

  独立的按 profile 网关(跨 UNIX 用户拆分的机群、或 HERMES_HOME 在 profiles/ 之外)需要 --force:
    hermes -p coder gateway install --force

  多路复用缺口关闭前的临时兼容路径:在 profiles/coder/config.yaml 设
  gateway.standalone: true,然后等宿主机网关重扫(<=30s)或发送其 rescan-profiles 控制动词。
  (gateway.standalone 是多路复用缺口修复期间的临时兼容垫片;
  缺口修复后会移除它——计划用 `hermes gateway migrate --multiplex` 折叠这个 profile。)

当宿主机网关已在跑且服务该 profile 时,第一行读 The host gateway already serves profile 'coder'.,带属主 PID 和服务集合,指针是 hermes -p default gateway restart。仪表盘对命名 profile 的 Start 按钮返回同样拒绝。

临时:gateway.standalone: true

临时向后兼容,不是我们保留的拓扑

只多路复用是方向:每宿主机一个网关服务每个 profile。切换落地时并非每个缺口都已关闭——次要 profile 上的 WhatsApp 桥接和 relay、以及仪表盘作用域是仍开着的;按 profile 停/启/重启已由停驻关闭——依赖按 profile 网关的机群一夜之间失去了它们。 gateway.standalone: true 的存在是为了让这些机群在那些缺口修复期间继续工作。缺口修复后会移除它,并提前在发行说明中通知;每个打印它的地方都这么说。不要在它上面建新设置:如果你从零开始,跑宿主机多路复用器。如果今天有个缺口挡你,设这个键,并为该缺口开 issue 或点赞,好让我们早点移除垫片。

在 profile 自己的 config.yaml 设 gateway.standalone: true:

# profiles/coder/config.yaml
gateway:
  standalone: true

然后宿主机网关不服务该 profile,启动日志记录 profile 'coder' is standalone (gateway.standalone: true); not served by this gateway。hermes -p coder gateway install|start|run 不用 --force 即可工作。如果运行中的宿主机仍在其服务集合里列出该 profile,命令在它重扫前拒绝:等下次重扫(正常运行下最多 30 秒),或向宿主机网关发 rescan-profiles 控制动词。无需重启宿主机。

移除该键让 profile 重新对宿主机合格。当该 profile 自己的网关在线时,宿主机跳过添加它并记录必须先停掉它。停掉那个网关;宿主机下次重扫接管该 profile。

独立 profile 的适配器、cron、webhook 入口和看板通知只在它自己的网关运行时跑,不在宿主机多路复用器或 hermes serve 下。把 webhook 客户端指向独立网关自己的监听器;宿主机的 /p/<profile>/ 入口不再服务它。那个监听器从该 profile 自己的 .env 或 config.yaml 解析端口,因此当它和宿主机网关都启用 API 服务器或 webhook 入口时,给该 profile 自己的 API_SERVER_PORT / WEBHOOK_PORT;两个网关都留默认会都试图绑定它们。cron 目标选择器仍把独立 profile 列为 bot-chat:<name> 目标,但宿主机无法投递到这些目标。

hermes -p coder gateway status 在该 profile 自己的网关状态前打印 standalone by config (gateway.standalone: true),hermes gateway status(默认)在服务集合后把它列为 standalone by config: coder。hermes gateway migrate --multiplex 不动该 profile,打印 Standalone by config (gateway.standalone: true), left alone。WhatsApp 桥接和 relay 在该 profile 自己的网关里跑,与任何独立网关一样。

--force 不是这个垫片的路径;它仍是拒绝命令点名的两种边界情况的逃生(跨 UNIX 用户拆分的机群、HERMES_HOME 在 profiles/ 之外):它装一个真正的按 profile 服务,那个服务(其 ExecStart 不带 --force)此后正常启动。

多路复用开启时变了什么

多路复用改变几样东西的行为。这些都不适用于用 gateway.standalone: true 退出、或在被阻塞宿主机上跑独立 --force 网关的 profile。

1. 次要 profile 不得启动自己的网关

多路复用器跑着时,命名 profile 的 gateway run attach 到它;gateway install 拒绝创建另一个进程(退出码 78)。CLI 在触碰服务管理器之前拒绝,防止永久失败的 systemd unit 或 launchd 重生循环。用上面的按 profile stop、start、restart 命令管理宿主机内的卫星。默认 profile 的 hermes gateway stop 仍把每个已服务 profile 下线。仪表盘和桌面应用跟随 CLI:对已服务 profile,"Stop" 停驻它、"Start" 取消停驻(见上);对未停驻的命名 profile,"Start" 以同样解释回 409(在 System 页渲染为内联提示)。"Restart" 重启多路复用器(真正服务该 profile 的进程),而非派生一个只会失败的 -p coder gateway restart。因为那次重连会重连设备上所有 bot,两个应用先问 "Restart the shared gateway? All bots on this device reconnect: default, coder, research"(列表是运行中网关的 served_profiles),完成后报 "Shared gateway restarted (3 bots)"。独立 profile 保持普通重启。/api/status?profile=coder 带与 gateway_shared_with 相同的列表(独立网关为 null)。"已服务"从运行中网关自己的记录(默认 home 的 gateway_state.json 里的 served_profiles)读取,因此即使多路复用器只通过默认 profile 环境里的 GATEWAY_MULTIPLEX_PROFILES 启用、或网关启动后才添加 profile,它也保持正确。

setup 流程遵循同一规则:hermes -p coder setup gateway、hermes -p coder setup、hermes -p coder gateway setup 和 hermes -p coder import 配置该 profile 的 bot,但为已服务 profile 跳过"安装网关后台服务"步骤,打印 "Profile 'coder' is already served by the default multiplexer",而非注册一个只会躺死的 stray unit 或 plist。加上 bot token,运行中的多路复用器会捡起它。

多路复用器是唯一入站进程;第二个 profile 网关会双重绑定该 profile 的平台。故意要独立进程的 profile 用 gateway.standalone: true 退出(见不再新建按 profile 网关);只在边界挡住折叠时传 --force(run、start、install、restart 都接受)。因此本页前面的跨 profile 生命周期包装脚本在多路复用模式下不用——直接管理宿主机或其命名 profile。

2. HTTP 入站平台通过 /p/<profile>/ URL 前缀访问

次要 profile 的 HTTP 入站流量到达默认 profile 的一个监听器,带 profile 前缀,不是第二个端口:

# 默认 profile
POST http://host:8644/webhooks/<route>
# "coder" profile,同一监听器
POST http://host:8644/p/coder/webhooks/<route>

前缀中未知或未配置的 profile 返回 404。共享监听器是默认 profile 的 api_server 端口(或未启用 API 服务器时的 webhook 端口);它服务三类带 profile 前缀的路径:

  • api_server 和 webhook 被镜像,绝不重复。 /p/coder/v1/... 和 /p/coder/webhooks/<route> 由默认 profile 自己的适配器在 coder 作用域下应答。因此次要 profile 必须不自己启用 api_server 或 webhook(仪表盘以 409 拒绝;次要 .env 里的 API_SERVER_KEY 或 WEBHOOK_ENABLED 只接好凭据而不启动监听器)。
  • 其他每个入站端口平台都跑在共享监听器模式。 配置了 Twilio SMS、LINE、Teams、BlueBubbles、Microsoft Graph、WhatsApp Cloud、WeCom 回调或飞书 webhook 模式的次要 profile,会得到一个不带端口构建的自己的适配器实例;默认监听器把 /p/<profile>/<该适配器惯常路径> 转发给它。见多路复用器下的入站端口平台。
  • WhatsApp(桥接)和 Relay 是默认 profile 拥有的共享入口。 多路复用器绝不为次要 profile 启动它们:profiles/work/.env 里的 WHATSAPP_ENABLED=true 自己什么都不做。在默认 profile 上启用并配置它们(其入站经 profile_routes 路由到各 profile),或在次要 profile 里禁用。网关为每个被跳过的次要平台记一行 INFO,如果没有 profile 跑它,则记一条 WARNING 说该平台未被服务;hermes gateway status --profile work 显示 whatsapp: not served under multiplex (shared ingress owned by default)。唯一例外是用 gateway.standalone: true 退出的 profile——它在自己的网关里跑自己的 WhatsApp 桥接和 relay,与任何独立网关一样。

认证跟随 URL 中命名的 profile。无前缀端点继续用默认监听器已有的凭据。

  • /p/coder/... 的 API 服务器请求必须用 ~/.hermes/profiles/coder/.env 里的 API_SERVER_KEY;默认监听器的 key 被拒。多路复用器下该 key 只认证前缀——它不在次要 profile 里开第二个 api_server 监听器,因此你不需要在次要 config.yaml 固定 platforms.api_server.enabled: false。
  • 一个指向 coder 的 webhook 路由必须在默认 profile config.yaml 里既有路由专属 secret 旁声明 profile: coder。该 secret 此后只在 /p/coder/webhooks/<route> 被接受,在其他每个 profile 前缀被拒。
  • 不带 profile 的 webhook 路由仍是默认 profile 路由,不能通过命名 profile 前缀访问。动态订阅同样绑定:hermes webhook subscribe <name> --route-profile coder 把 profile: coder 写进默认网关的 webhook_subscriptions.json 并打印 /p/coder/webhooks/<name> URL(hermes webhook ls 显示绑定)。用 --route-profile,别用全局 -p coder:-p 会把订阅写进 coder 自己的订阅文件,而默认网关的 webhook 适配器从不读它。
  • 投递遵循同一绑定。profile: coder 路由的回复(或 deliver_only 消息)经 coder 的该 deliver 平台适配器发出,deliver_extra.chat_id 未设时回退到 coder 的主频道,github_comment 投递用 profiles/coder/.env 里的 GH_TOKEN / GITHUB_TOKEN 跑 gh。如果 coder 没有该平台适配器,投递失败(502),而不是作为另一个 profile 的 bot 发帖;默认路由同样绝不借用只在次要 profile 上启用的平台。
  • /p/coder/api/platforms/<platform>/events 回调由 coder 的适配器校验并派发;coder 没有则回调是 503。

命名 API 请求在目标 profile 没有 API_SERVER_KEY 时关闭失败。安全配置错误仍致命:例如一个 open 自有策略平台没有 GATEWAY_ALLOW_ALL_USERS 或其平台专属 allow-all 选项,仍中止网关启动,而非静默丢弃该不安全 profile。

多路复用器下的入站端口平台

独立的 hermes -p coder gateway run 把 coder 的 Twilio、LINE、Teams 等 webhook 服务器绑在自己的端口上。多路复用器下这些适配器仍是 coder 的——profiles/coder/.env 里同样的凭据、同样的 config.yaml、回复经 coder 的频道发出——但它们不绑端口。默认 profile 的共享监听器把 /p/coder/<path> 转发给它们,<path> 正是该适配器独立服务时的路径。请求由 coder 的适配器用 coder 的 secret 校验(Twilio auth token、LINE channel secret、Teams 应用凭据、BlueBubbles 密码……),并在 coder 的运行时作用域下运行;默认 profile 自己的 /path 不动,没有该路径适配器的 profile 拿到 404,绝不会是另一个 profile 的 bot。

平台共享监听器上次要 profile 的回调 URL用命名 profile 的什么校验
Twilio SMS(sms)https://<host>/p/<profile>/webhooks/twilioTWILIO_AUTH_TOKEN 签名(SMS_WEBHOOK_URL 必须是此 URL)
LINE(line)https://<host>/p/<profile>/line/webhook(媒体:/p/<profile>/line/media/...)LINE_CHANNEL_SECRET
Microsoft Teams(teams)https://<host>/p/<profile>/api/messagesTEAMS_CLIENT_ID 的 Bot Framework token
BlueBubbles(bluebubbles)http://<host>/p/<profile>/bluebubbles-webhook(自动向服务器注册)BLUEBUBBLES_PASSWORD
Microsoft Graph(msgraph_webhook)https://<host>/p/<profile>/msgraph/webhookextra.client_state
WhatsApp Cloud(whatsapp_cloud)https://<host>/p/<profile>/whatsapp/webhookWHATSAPP_CLOUD_APP_SECRET / verify token
WeCom 回调(wecom_callback)https://<host>/p/<profile>/wecom/callback应用的回调 token / AES key
飞书 webhook 模式(feishu)https://<host>/p/<profile>/feishu/webhookFEISHU_VERIFICATION_TOKEN / FEISHU_ENCRYPT_KEY

<host> 是默认 profile 监听器前面的公网主机名(隧道、反向代理);profile 配置里自定义的 webhook_path 把路径相应移到 /p/<profile> 之后。网关启动时打印确切 URL:

[sms] profile 'coder' is served on the default profile's shared listener:
http://127.0.0.1:8642/p/coder/webhooks/twilio (point the vendor's callback URL at this path ...)

每个状态界面都重复它,因此你知道往厂商控制台粘贴什么:

$ hermes -p coder gateway status
✓ Gateway is running via the default-profile multiplexer
  Manage it from the default profile: hermes gateway status

Inbound callback URLs on the shared listener:
  line: http://127.0.0.1:8642/p/coder/line/webhook
  sms: http://127.0.0.1:8642/p/coder/webhooks/twilio

默认 profile 上的 hermes gateway status 和 hermes status 按已服务 profile 列出同样的 URL,仪表盘 Channels 页和桌面 Messaging 页在查看该 profile 时把它们显示为各平台的 ingress_url。默认自己的 api_server 和 webhook 对已服务 profile 也同样报告为 connected,ingress_url 为 http://127.0.0.1:8642/p/coder/v1(分别 .../p/coder/webhooks/<route>)——因为该 profile 没有自己的适配器;是默认监听器在 /p/coder/ 前缀下应答。次要 .env 里的按 profile SMS_WEBHOOK_PORT、LINE_PORT、TEAMS_PORT 等在多路复用器下被忽略(什么都不绑);该 profile 一旦跑自己的独立网关就重新生效。

3. 按凭据的平台仍需每 profile 自己的 token

轮询/连接平台(Telegram、Discord、Slack、Matrix、Signal……)多路复用下工作良好,但每个启用它的 profile 必须提供自己的 bot token——同一 token 不能被两个 profile 同时轮询。如果两个 profile 配置了相同的 (platform, token),网关记一条错误,点名两个 profile,并停驻重复适配器(运行时状态显示为 fatal / duplicate_credential),而首个认领者和其他每个 profile 继续跑——网关本身不退出。默认 profile 的适配器先连接并认领凭据,因此被停驻的总是次要的适配器(见Token 冲突安全——规则不变,只是现在在一个进程内强制)。

4. 会话键按 profile 命名空间

每个 profile 的会话在 agent:<profile>:… 命名空间下,因此同一平台/聊天上的两个 profile 在共享会话存储里绝不冲突。默认 profile 逐字节保留历史的 agent:main:… 命名空间,因此已有默认 profile 会话不受影响——无迁移、无孤儿历史。每个读回键的网关路径——重启后的委派完成、关闭通知、对兄弟运行的按用户线程 /stop、/undo、QQ 审批按钮——也接受 agent:<profile>:… 形态,因此次要 profile 获得与默认相同的行为。唯一一个会与默认命名空间冲突的 profile 名——一个真叫 main 的 profile——以 agent:main~:… 为键,因此它保留自己的会话和自己的 profiles/main/state.db。

每个 profile 的行落在它自己的 state.db:命名 profile 在 profiles/<name>/state.db,默认 profile 在启动 home——即使写入发生在另一个 profile 的路由轮或后台 tick 内。桌面/TUI 后端自己的存储同样固定在它启动时的 home,Bot Chat 的子 agent(prompt.background)持久化在其父会话旁。

5. 一个 PID/锁和一个状态界面

只有一个进程级 PID 和锁(多路复用器,在默认 home 下)。默认 profile 上的 hermes status 报告多路复用器并列出它服务的 profile(Serves: coder, research)。hermes -p coder status 和 hermes -p coder gateway status 报告"running via the default-profile multiplexer"而非"stopped"。仪表盘的 /api/status?profile=coder / Channels 页把多路复用器报告为 coder 的运行网关,coder 自己的适配器作为其平台。单个 gateway_state.json 在默认 home 下:次要适配器在那里作为 <profile>:<platform> 条目出现在 served_profiles 旁;不写按 profile 网关状态文件。

hermes -p coder cron status 点名单一宿主机网关和它服务的 profile——Scheduler host: the host gateway (PID 4211) serving profiles default, coder——然后检查 coder 自己的 ticker 心跳和上次成功 tick。缺失或陈旧心跳产生警告,而非无条件运行判定。cron list 和 cron create 在已服务 profile 没有新鲜心跳时也警告。cron status 补充那些轻量检查不读的 tick 失败详情。

没有网关拥有宿主机角色时,cron status 告诉你启动那个宿主机网关(hermes --profile default gateway install / gateway run)并确认它服务该 profile。安装按 profile 服务只在 LEGACY (pre-multiplex topology, not recommended) 下显示——它会在宿主机上启动第二个网关进程。hermes doctor 遵循同一规则——s6 下它报 Host gateway: the host gateway (PID 4211) serving profiles default, coder 而非按 profile 槽位数,把仍受监管的按 profile 槽位标记为 LEGACY,即使你从已服务 profile 跑 doctor 也检查宿主机 systemd unit 的 linger。state.db 持有者行也点名共享宿主机进程,因此"3 process(es) holding the DB open"说明停掉它会影响哪个网关和哪些 profile。

什么不变

按 profile 的 .env 凭据隔离保留,甚至更严:一个 profile 的 key 从自己作用域解析,绝不并入共享环境。MCP 服务器、看板 worker 等子进程只看得到自己 profile 的机密——包括外部机密源(1Password、Bitwarden 等)注入的凭据:为 profile B 启动的 stdio MCP 服务器收到 B 对该名字的值,B 没有则什么都没有,绝不是默认 profile 的。MCP 服务器按 profile 连接:两个 profile 都用自己的 token 命名一个叫 github 的服务器,得到两条连接,各自只看自己的工具;mcp_servers 条目完全相同(同一路由和凭据,含 mTLS client_cert/client_key)的 profile 共享一条连接,属主的 /reload-mcp 重新登记共享 profile 的工具而无需它们重载。auth: oauth 服务器绝不跨 profile 共享:每个 profile 在自己的 mcp-tokens/ 下持有自己的 token 并开自己的连接。启动时一个接一个连 profile,在一个 profile 内最多同时 mcp.discovery_concurrency 个服务器(默认 4,0 = 不限),因此带许多 stdio 服务器的 profile 机群不再在同一瞬间派生每个 helper 进程。信任策略按 profile 保留:一个 trust: untrusted 共享 trust: full profile 连接的 profile,每次可写调用前仍被询问,supports_parallel_tool_calls 只对设置它的 profile 生效。终端设置(terminal.backend、terminal.cwd、terminal.docker_volumes、terminal.docker_shared_container_key、SSH 目标等)同样每轮路由按 profile 解析:省略某个终端键的 profile 拿到文档默认值,绝不拿启动 profile 的值;config.yaml/.env 无法解析的 profile 拒绝终端执行,而非在另一个 profile 的沙箱策略下运行。媒体投递凭据守卫(MEDIA: 附件背后的拒绝列表——.env、auth.json、config.yaml、state.db、会话转录、OAuth token 存储)覆盖 profiles/ 下每个 profile,因此任何 profile 的轮次都不能把另一个 profile 的机密或聊天历史附到回复上。授权也按 profile:GATEWAY_ALLOW_ALL_USERS、GATEWAY_ALLOWED_USERS 和每个平台允许列表或 allow-all 选项都从属主 profile 的 .env 读取——默认 profile 选开放访问绝不打开次要 profile 的 bot,只在自己 .env 里选开放的次要 profile 也被尊重。写在 profile config.yaml 里的按 bot 行为(require_mention、mention_patterns、allow_bots、reactions、auto_thread、free_response_auto_thread、dm_policy、ignored_channels、Matrix session_scope 等)同理:次要 profile 的 YAML 绝不落入共享进程环境,因此不能成为默认 profile 的策略,默认 profile 的 YAML 也绝不治理次要 bot。terminal.env_passthrough 允许列表、元宝自动指定的主频道、保护每个 profile 自己 config.yaml 的写入守卫也按 profile 解析。看板、按 profile 的 skill/记忆/SOUL 和模型路由都按 profile 行为,与独立网关时完全一样。

出站身份也按 profile。为 profile P 运行的一轮调用 send_message 工具(发送、回应、媒体)时,经 P 自己的 bot 发帖;P 会话的"Gateway shutting down/restarted"和 /update 通知、从 P 聊天里设的 /loop 唤醒、以及 P 的 Discord bot 的未授权斜杠操作者告警(发到 P 的主频道)也一样。如果 P 没有该平台的已连 bot,发送以明确错误失败——绝不回退到默认 profile 的 bot。

工具和记忆提供商凭据遵循同一规则。托管 OCR(FIRECRAWL_API_KEY)、Modal / Browser Use 云门、mem0 OSS OpenAI key、xAI 视频、以及每个记忆提供商身份(MEM0_USER_ID、SUPERMEMORY_CONTAINER_TAG、RETAINDB_PROJECT、OPENVIKING_ACCOUNT/USER、HINDSIGHT_BANK_ID、HERMES_HONCHO_HOST)都从被路由 profile 的 .env 读取,因此次要 profile 的记忆落到它的账号/bank/项目(或提供商的按 profile 默认),绝不落默认 profile 的。自定义端点随其 key 一起走——OPENAI_BASE_URL、XAI_BASE_URL、NOUS_INFERENCE_BASE_URL、GATEWAY_PROXY_URL、Firecrawl / Browserbase / RetainDB / Supermemory / Honcho / Hindsight URL——因此一个 profile 的 key 绝不发到另一个 profile 的代理或自托管服务器。WEIXIN_HOME_CHANNEL、HERMES_LANGUAGE 和 display.language、hooks.outbound[].secret_env 同样按 profile,被逐出次要会话的会话末记忆抽取在该 profile 作用域下跑。

每轮运行时设置也跟随被路由 profile:agent.max_turns、fallback_providers、file_read_max_chars、tool_output.*、browser.* 超时、timezone(包括交给 execute_code 沙箱的 TZ)、媒体投递策略(gateway.strict、media_delivery_allow_dirs、trust_recent_files*)以及辅助调用用的 Nous auth.json,都从服务该轮的 profile 读取,绝不从网关启动时的 profile 读。按 profile 状态文件(processes.json、checkpoints/、沙箱快照存储、飞书评论规则/配对)和网关钩子同理:每个 profile 的 hooks/ 目录单独加载,只为该 profile 的事件触发。shell 钩子以被路由 profile 的 HERMES_HOME 运行,环境里不带默认 profile 的机密,其 stdin 载荷带一个 profile 字段,点名触发它的 profile。

按 profile 隔离什么

速查表:多路复用的一轮从自己的 profile 解析什么,绝不与默认或任何兄弟共享:

关注点从哪里解析profile 缺它时的行为
提供商 key、bot token、config.yaml 里的 ${VAR} 引用该 profile 自己的 .env(其机密作用域)未解析 / 无适配器——绝不是默认 profile 的值
授权(GATEWAY_ALLOW_ALL_USERS、GATEWAY_ALLOWED_USERS、按平台允许列表和 allow-all 选项)属主 profile 的 .env 和 config.yaml关闭——默认 profile 的开放绝不打开次要 bot
斜杠命令门控(allow_admin_from、user_allowed_commands、group_allow_admin_from;见斜杠命令)收到消息的那个 bot 的 profile——次要自己的平台 extra 块治理其 bot,不是默认 profile 的关闭失败:多路复用器还没加载其配置的已服务 profile 以空管理员列表和无用户启用命令门控,只有永远允许的底线(/help、/whoami)运行——绝不是默认 profile 的开放策略
HTTP 端点(/p/<profile>/api/...、/p/<profile>/webhooks/...、平台事件回调)命名 profile 的 API_SERVER_KEY、profile: 绑定的 webhook 路由、其自己的适配器401/404;无适配器的投递是 502/503,绝不是另一个 profile 的 bot
入站端口平台(/p/<profile>/webhooks/twilio、/p/<profile>/line/webhook、/p/<profile>/api/messages……)命名 profile 自己的适配器及其 secret(Twilio auth token、LINE channel secret、Teams 应用、BlueBubbles 密码……);回复经该适配器发出secret 错误时 401/403,profile 无此适配器时 404——绝不是默认 profile 的适配器
适配器设置(*_REQUIRE_MENTION、*_REACTIONS、*_ALLOW_BOTS、*_PROXY、Discord allow_mentions、Matrix allowed_users / ignore_user_patterns、webhook 主机/端口/URL、Matrix 线程/会话/E2EE 策略、Discord 回填/附件上限、Buzz 回复模式、A2A agent 卡片 / 公网 URL、WhatsApp 桥接策略、元宝主频道)属主 profile,按此顺序:显式 .env 值 → 其 config.yaml → 适配器默认适配器文档默认——绝不是默认 profile 的设置。单 profile 安装保持 env 覆盖 YAML,与各平台页文档一致
MEDIA: 附件拒绝列表检查时枚举 profiles/ 下每个 home 加默认 home一轮绝不能附另一个 profile 的 .env、auth.json、state.db、会话或 token 存储
stdio MCP 子进程环境安全基线 + 该 profile 作用域内的机密源名字 + 服务器自己的 env:profile 缺的名字不在子进程里——无默认 profile 透传
出站 egress(send_message、关闭/重启//update 通知、/loop 唤醒、profile: 绑定的 webhook 投递、github_comment token)该 profile 自己已连的适配器和 .env明确失败;绝不经默认 profile 的 bot 发帖
会话命名空间agent:<profile>:…(默认保留 agent:main:…)同聊天上两个 profile 绝不共享历史
日志该 profile 自己 home 下的 agent.log / errors.log / gateway.log—
终端沙箱设置(terminal.*、SSH 目标)该 profile 的 config.yaml文档默认;配置不可解析 → 拒绝执行
一轮的工作目录(未设 terminal.cwd)与独立网关同规则:本地后端 $HOME,否则沙箱默认绝不是多路复用器进程启动时的目录
命令审批(command_allowlist、"always" 选择)该 profile 自己的 config.yaml默认 profile 的"always"绝不预批次要 profile 的命令;次要的选择存进自己配置
沙箱凭据文件挂载(terminal.credential_files)、security.redact_secrets、browser.* 引擎/有头标志、lsp.*、辅助提供商健康标记、logs/mcp-stderr.log该 profile 自己的 config.yaml / .env文档默认——绝不是启动 profile 的缓存值
云 SDK 凭据客户端(Bedrock boto3 客户端 + 模型发现、Azure Entra 凭据)、凭据拉取的目录(DeepInfra、Copilot 上下文限制、Nous 推理上限、Ramp Router efforts、xAI / OpenRouter 图像模型、自定义端点 /models)、Camofox VNC 地址、computer-use 辅助视觉路由、skill 同步推送、远程后端探测文本、学到的图像 token 成本、display.skin、访客 mint 退避、横幅 skill、元宝 "active" 适配器、Langfuse 客户端该 profile 自己的 .env / config.yaml / <home>/cache文档默认——绝不是启动 profile 的缓存值或其凭据
会话搜索旋钮(sessions.cjk_fts、sessions.search_slow_ms)该 profile 的 config.yaml文档默认——绝不是默认 profile 的桥接值
RoomLink 能力目录及其向远程 Bot 宣告的签名执行策略(approvals.mode、agent.max_turns、platform_toolsets.api_server)请求命名的已服务 profile(/p/<profile>/v1/room-members/...、RPC profile 参数);每个目录必须有 target_profile——无 HERMES_PROFILE 回退邀请/能力失败时点名违规的 target_profile;不存在的 profile 被拒,绝不从启动 profile 配置解析
平台代理(TELEGRAM_PROXY、DISCORD_PROXY、HTTPS_PROXY……)该 profile 自己的 .env直连——绝不是默认 profile 的代理
桌面/仪表盘后端的 MCP 发现每个已服务 profile home 一次在另一个已建 agent 的 profile 之后选中的 profile 仍发现自己的 mcp_servers
从桌面 / TUI 会话改的设置(/busy、/verbose、/approval、/cwd、主题和显示开关)拥有会话的那个 profile 的 config.yaml,即使 RPC 只带会话 id写会话自己的 profile;启动 profile 的 config.yaml 及其 TERMINAL_CWD 绝不碰
桌面/仪表盘后端的 MCP 连接和按 profile cron ticker即使 gateway.multiplex_profiles 关闭也按已服务 profile 为键——与多路复用器同规则其他凭据的同名 mcp_servers 条目是自己的连接;已服务 profile 绝不以另一个 profile 身份调服务器
仪表盘动作(桌面/仪表盘派生的 hermes -p <name> …)固定到该 profile HERMES_HOME 的清洗后子进程环境子进程加载自己的 .env;仪表盘 profile 的 token 和端口不被继承
为已服务 profile 行事的每个子进程(斜杠 worker、Bot Chat 投递、A2A 转发、key_cmd helper、浏览器驱动)该 profile 自己的 .env + 机密源,跑在凭据清洗后的基线上——无论有无 gateway.multiplex_profiles(桌面/仪表盘 ?profile= 路由也算)子进程里没有——只经 systemd / Compose / shell 到达启动进程的 key 绝不被另一个 profile 的子进程继承
为另一个 profile 派生的子进程里的授权门(*_ALLOWED_USERS / *_ALLOWED_CHANNELS / *_IGNORED_CHANNELS / *_ALLOW_ALL_USERS / *_ALLOW_BOTS、GATEWAY_ALLOW*)——仪表盘 hermes -p <name> 动作、看板 worker、Bot Chat 投递、更新后按 profile 的 gateway restart子进程自己的 .env / config.yaml,由子进程自己加载关闭(适配器文档默认)——unit 文件或 shell 导出到派生进程的门在子进程启动前丢弃,因此 profile B 绝不强制 profile A 的频道或用户列表;同 profile 子进程保留它
把已服务 profile 镜像进活跃 HERMES_HOME 环境变量供旧读者(Hermes WebUI)的嵌入宿主机里的路由 profile 检测宿主机用 hermes_constants.pin_process_hermes_home() 固定的启动 home;MCP 连接键、已服务 profile 子进程的启动环境剥离、桥接的 allow-all 种子和 terminal.* 环境桥守卫都对照它比较没固定时活跃环境变量就是启动 home,与以前完全一样——从不改 HERMES_HOME 的宿主机什么都不需要
同时服务另一个 profile 的 hermes serve / 仪表盘进程中,启动(默认)profile 自己的凭据其 .env + 机密源,跑在第一个其他 profile 被服务那一刻冻结的进程环境上;之后不再重读只在进程环境里轮换的凭据(systemctl set-environment、未重新 exec 的刷新 op run 包装)要等进程重启才被捡起——把轮换的 key 放 .env 或机密源,或轮换后重启
cron .env 调优(HERMES_CRON_TIMEOUT、HERMES_MODEL 回退、HERMES_CRON_MAX_PARALLEL、预填文件)、worker / Bot Chat 子进程环境该 profile 自己的 .env;子进程绝不继承默认 profile 的 .env 设置或桥接的 TERMINAL_* 策略cron 默认 / 模型拒绝,与独立 hermes -p <name> gateway run 完全一样
一个 profile 任务的看板 worker 和通知被指派者的 .env + config.yaml(工具集固定、终端后端、媒体策略、显示语言)—
/loop tick、background_process_notifications 门、notice_delivery、后台进程检查点恢复属主 profile 的 state.db / config.yaml / processes.json—

刻意共享的:进程、其 PID/锁和 gateway_state.json(默认 home)、那一个 HTTP 监听器、以及 profile_routes 表(声明在默认 profile 上)。

哪些 profile 被服务

gateway.multiplex_profiles: true 服务默认 profile 加 profiles/ 下每个活跃命名 profile,除了那些用自己 config.yaml 里 gateway.standalone: true 退出的——独立 profile 跑自己的网关,不被宿主机枚举(见不再新建按 profile 网关)。( former gateway.multiplex_profile_allowlist 键已退役;配置迁移把它从 config.yaml 移除,你不想服务又没退出的 profile 应改为归档或删除——hermes profile delete <name>,或把目录移出 profiles/。)被删 profile 留下墓碑,永不枚举;目录没了的 profile 绝不会被已服务轮次、cron ticker 或日志路由重建。

服务集合控制 /p/<profile>/ API 和 webhook 前缀、运行时状态、profile 路由合格性,以及进程内 cron 调度器 tick 哪些 profile(桌面后端的 ticker 每个周期重枚举同一集合——桌面运行时创建或删除的 profile 加入或离开 tick 集合而无需重启——并对运行中多路复用器或其自己网关已服务的 profile 退让)。作为 hermes -p <name> gateway run 启动的多路复用器也总是 tick 它自己 profile 的 cron 存储。

服务集合是活的。多路复用器运行时创建的 profile(hermes profile create、仪表盘、桌面或 TUI)立即被服务:创建者经控制 socket ping 多路复用器,多路复用器还每 30 秒重扫 profiles/ 作为安全网。新 profile 的适配器在其 config.yaml/.env 带 bot token 那一刻构建(创建者通常先创建再加 token),默认 profile gateway_state.json 里的 served_profiles 更新,hermes -p <name> gateway status 报告它已服务——无需重启,其他 profile 的适配器和在途轮次不动。删除 profile 同样停并解路由其适配器,hermes profile rename 在目录移动前解路由旧名、热服务新名(旧名不会被仍绑定它的适配器或 cron ticker 复活)。一凭据一轮询者规则仍适用:热添加的 profile 复用另一个 profile 的 token 会以 duplicate_credential 错误停驻,绝不作为第二个轮询器启动。

把共享 bot 聊天路由到 profile(profile_routes)

多路复用器按凭据(每个 profile 自己的 bot token)或按URL 前缀(HTTP 平台的 /p/<profile>/)选 profile。当多个社区共享一个 bot token——例如一个 Discord bot 服务多个 guild——你还可以用 gateway.profile_routes 把特定用户/guild/频道/线程路由到不同 profile:

gateway:
  multiplex_profiles: true
  profile_routes:
    # 整个 Discord 服务器 → 一个 profile
    - name: acme-server
      platform: discord
      guild_id: "1234567890"
      profile: acme

    # 该服务器里一个频道 → 另一个 profile
    - name: acme-support
      platform: discord
      guild_id: "1234567890"
      chat_id: "9876543210"
      profile: acme-support

    # 一个 Telegram 群组(无 guild 概念——仅 chat_id)
    - name: tg-group
      platform: telegram
      chat_id: "-1001234567890"
      profile: tg-profile

    # 一个 WhatsApp DM——写电话号码;JID 和 LID 形式也匹配
    - name: owner-whatsapp
      platform: whatsapp
      chat_id: "15551234567"
      profile: owner

    # 一个跨 DM、群组和频道的 Teams 用户(精确发送者 id)
    - name: teams-owner
      platform: teams
      user_id: "00000000-0000-0000-0000-000000000000"
      profile: owner

路由按累加特异性匹配:user_id = 16、thread_id = 8、chat_id = 4、guild_id = 2。因此 user_id + chat_id(20)压过单独 user_id(16),后者压过每个仅位置路由(最多 14)。所有声明字段必须同时成立(AND),同分保持声明顺序,以频道为键的路由也匹配父频道是该频道的线程/论坛帖。不匹配任何路由的消息留在默认/活跃 profile。被路由 profile 获得上文描述的完整按 profile 隔离(配置、skill、记忆、凭据、会话命名空间)。路由在每个平台适配器上工作,不只是 Discord。

user_id 是入站消息的发送者,精确相等比较。它只和报告它的适配器一样可信,因此只在入口认证发送者的平台上把它当授权输入。某些平台上发送者 id 还按租户命名空间——Slack 用户 id 是工作区本地——因此在一个服务多个工作区或服务器的网关上,把 user_id 与该作用域的 guild_id(Discord guild、Slack 工作区、Matrix 服务器)配对,而非单靠 id。

省略 user_id 为向后兼容保持路由不受发送者约束。把它设为 null、空字符串或空白会使该路由无效,而非放宽到平台上每个发送者。

发送者路由选一个 profile;它不是默认拒绝授权。不匹配任何路由的发送者落到默认/活跃 profile,与未路由频道完全一样。要给一个人特权 profile、给其他人受限 profile,先声明特权发送者路由,其后加一条平台级 catch-all 到受限 profile,并保持平台自己的入口允许列表到位。

路由只作用于默认 profile 的 bot 收到的消息,除非用 bot_profile: <profile> 点名另一个 bot。Telegram DM 对每个 bot 用同一 chat_id(用户 id),因此没有这个,一个为共享 bot 设的 chat_id 路由也会捕获该用户与次要 profile 专用 bot 的 DM。到达次要 profile 自己 bot 的消息留在该 profile:

    # 把一个用户与 team_b 自己 bot 的 DM 钉到第三个 profile
    - name: teamb-owner-dm
      platform: telegram
      bot_profile: team_b
      chat_id: "72719239"
      profile: ops-for-team-b

被路由消息的授权永远由接收 bot 的 profile(其 token 和允许列表)决定,包括 agent 忙碌时发送的后续和轮次中检查如 /topic 或 /stop;被路由 profile 自己不需要允许列表副本。没有自己 bot 的被路由 profile 在网关重启后也经共享 bot 收到后台通知(进程完成、心跳、异步委派结果)。

在 WhatsApp 和 WhatsApp Cloud 上,chat_id 路由跨用户身份形式匹配:裸电话号码(15551234567)、JID(15551234567@s.whatsapp.net)和 LID(…@lid)一旦桥接配对后都指同一人(同样的规范化会话键和适配器允许列表已在用)。你可以把电话号码放 profile_routes,入站 DM 无论 WhatsApp 投递 JID 还是 LID 都匹配。还没有 LID 映射时,号码形式仍匹配 JID(剥掉后缀),但无法解析未知 LID——那条入站落到默认 profile 直到映射出现。群聊(…@g.us)不是发送者身份,仍精确匹配。Telegram 数字 id 不变。

profile_routes 需要 gateway.multiplex_profiles: true;多路复用关闭时路由被忽略。如果显式路由匹配但其目标 profile 未安装(或被删),网关拒绝该入口并记录路由和目标。它不跑默认 profile。不匹配路由的流量保持历史默认 profile 行为。

被路由 profile 拥有的 cron 任务也经共享 bot 投递,但只投到一条带 chat_id/thread_id 的启用路由映射到该 profile 的目标(guild_id + chat_id 路由使其频道合格)——被路由 profile 的任务投到未路由聊天(或路由到另一个 profile 的聊天)绝不经共享 bot 发送。仅 guild 的路由不使 cron 目标合格;为投递频道加一条 chat_id 路由。声明 user_id 的路由也不合格:cron 没有已认证的入站发送者,因此它们需要单独的仅位置路由。被路由 profile 为此不需要自己的 platforms.<platform> 块:共享 bot 的授权来自路由,而非卫星的配置。

一次启动、停止或重启所有网关

CLI 自带单 profile 生命周期命令。要跨每个 profile 行动,把它们包在 shell 循环里。把下面片段放进 ~/.local/bin/hermes-gateways 并 chmod +x:

#!/bin/sh
set -eu

# 创建/删除 profile 时在这里加或删 profile 名。
profiles="default coder personal-bot research"

usage() {
  echo "Usage: hermes-gateways {start|stop|restart|status|list}"
}

run_for_profile() {
  profile="$1"
  action="$2"
  if [ "$profile" = "default" ]; then
    hermes gateway "$action"
  else
    hermes -p "$profile" gateway "$action"
  fi
}

action="${1:-}"
case "$action" in
  start|stop|restart|status)
    for profile in $profiles; do
      echo "==> $action $profile"
      run_for_profile "$profile" "$action"
    done
    ;;
  list)
    hermes gateway list
    ;;
  *)
    usage
    exit 2
    ;;
esac

然后:

hermes-gateways start      # 启动每个已配置 profile
hermes-gateways stop       # 停止每个已配置 profile
hermes-gateways restart    # 重启全部
hermes-gateways status     # 跨全部状态
hermes-gateways list       # 委托给 `hermes gateway list`
TIP

default profile 用 hermes gateway <action>(无 -p)指向,不是 hermes -p default gateway <action>。上面的包装器处理两种形式。

管理一个 profile

每个 profile 安装的快捷命令:

coder gateway run        # 前台(Ctrl-C 停止)
coder gateway start      # 启动托管服务
coder gateway stop       # 停止托管服务
coder gateway restart    # 重启
coder gateway status     # 状态
coder gateway install    # 创建 LaunchAgent / systemd unit
coder gateway uninstall  # 移除服务文件

这些等价于 hermes -p coder gateway <action>——当 profile 别名不在 PATH 上、或你从脚本动态指向 profile 时有用。

服务文件

每个 profile 安装自己的服务,名字唯一,因此安装绝不冲突:

平台路径
macOS~/Library/LaunchAgents/ai.hermes.gateway-<profile>.plist
Linux~/.config/systemd/user/hermes-gateway-<profile>.service

默认 profile 保留历史名字:ai.hermes.gateway.plist / hermes-gateway.service。

查看日志

每个 profile 写自己的日志文件:

# 默认 profile
tail -f ~/.hermes/logs/gateway.log
tail -f ~/.hermes/logs/gateway.error.log

# 命名 profile
tail -f ~/.hermes/profiles/<name>/logs/gateway.log
tail -f ~/.hermes/profiles/<name>/logs/gateway.error.log

同时流式查看每个 profile 的日志:

tail -f ~/.hermes/logs/gateway.log ~/.hermes/profiles/*/logs/gateway.log

CLI 还有结构化日志查看器:

hermes logs -f                  # 跟随默认 profile
hermes -p coder logs -f         # 跟随一个 profile
hermes logs --help              # 过滤器、级别、JSON 输出

识别实际在跑什么

hermes profile list             # profiles + 模型 + 网关状态
hermes-gateways status          # 跨每个 profile 完整状态
launchctl list | grep hermes    # macOS——PID 和标签
systemctl --user list-units 'hermes-gateway-*'   # Linux——units

编辑配置

每个 profile 把配置放在自己目录里:

~/.hermes/profiles/<name>/
├── .env              # API key、bot token(chmod 600)
├── config.yaml       # 模型、提供商、工具集、网关设置
└── SOUL.md           # 人格 / 系统提示词

默认 profile 直接用 ~/.hermes/,同样三个文件。

用任意编辑器或 CLI 编辑:

hermes config set model.model anthropic/claude-sonnet-4    # 默认 profile
coder config set model.model openai/gpt-5                  # 命名 profile

编辑 .env 或 config.yaml 后,重启受影响网关:

coder gateway restart
# 或,全部:
hermes-gateways restart

让主机保持唤醒

网关进程可以跑一整天,但操作系统空闲时仍会尝试休眠。两种模式:

macOS——caffeinate

caffeinate 内置于 macOS,运行时阻止休眠。无需安装。

caffeinate -dis                    # 阻止显示器、空闲和系统休眠
caffeinate -dis -t 28800           # 同上,8 小时后自动退出
caffeinate -i -w $(cat ~/.hermes/gateway.pid) &   # 默认网关跑着时保持唤醒

# 持久:后台运行并忘掉
nohup caffeinate -dis >/dev/null 2>&1 &
disown

# 检查 / 停止
pmset -g assertions | grep -iE 'caffeinate|prevent|user is active'
pkill caffeinate
标志作用
-d阻止显示器休眠
-i阻止空闲系统休眠(默认)
-m阻止磁盘休眠
-s阻止系统休眠(仅接电 Mac)
-u模拟用户活动(防止锁屏)
-t NN 秒后自动退出
-w PPID P 退出时退出
合盖仍会休眠 Mac

caffeinate 无法覆盖 MacBook 上硬件驱动的合盖休眠。要合盖运行,改节能器/电池偏好或用第三方工具。

Linux——systemd-inhibit 或 loginctl

# 命令运行时抑制挂起
systemd-inhibit --what=idle:sleep --who=hermes --why="gateways running" \
  sleep infinity &

# 允许用户服务登出后继续跑(推荐)
sudo loginctl enable-linger "$USER"

启用 linger 后,你的 systemd 用户 unit(包括 hermes-gateway-<profile>.service)跨 SSH 断开和重启继续跑。

Token 冲突安全

每个 profile 每个平台必须用唯一 bot token。如果两个 profile 共享同一个 Telegram、Discord、Slack、WhatsApp 或 Signal token,第二个网关以点名冲突 profile 的错误拒绝启动。多路复用下,同规则只停驻重复 profile 的适配器,共享网关继续跑。

审计:

grep -H 'TELEGRAM_BOT_TOKEN\|DISCORD_BOT_TOKEN' \
     ~/.hermes/.env ~/.hermes/profiles/*/.env

从按 profile 网关迁移

如果你的 profile 今天各跑各的网关(每 profile 一个 systemd unit 或 launchd agent,来自多路复用之前的版本),默认网关的启动预检在折叠它们之前保持独立(未设置的默认绝不会双重绑定运行中的机群)。hermes update 替你折叠,除非真实边界挡住(见下);同一次折叠是一条命令,在半迁移宿主机上重跑(标志开着、落下一个 unit、两者之间崩溃)会完成工作,而非报告"already multiplexed":

hermes gateway migrate --multiplex --dry-run   # 打印计划和任何阻塞;不改任何东西
hermes gateway migrate --multiplex             # 应用(TTY 上要求确认;-y 跳过)

没有 --standalone 反向命令:按 profile 机群不是受支持目标。被阻塞的机群照原样继续跑,每个 profile 以 hermes -p <name> gateway install --force 为路径——或用 gateway.standalone: true 退出宿主机网关(见不再新建按 profile 网关),hermes gateway migrate --multiplex 尊重它。

Docker / Hermes Cloud(s6 监管容器)

官方镜像内每个 profile 有一个 s6 槽位(/run/service/gateway-<profile>)。容器启动把每个命名槽位注册为 down,并把其自启意图折叠进根槽位,因此全新启动已经多路复用。原地更新也不再需要容器重启来收敛:hermes gateway migrate --multiplex(以及 hermes update 跑的钩子)停驻任何仍起的命名槽位(s6-svc -d 加一个 down 文件,这样 supervisor 重启不会复活它),按启动用的同规则把其意图折叠进根槽位,并重启根槽位。注册为 down 的槽位从不阻塞——只有真正起来的槽位才阻塞。该命令从内部仍做不到的一件事是创建启动从未注册的根槽位;那种情况自己点名并要求容器重启。

hermes update 做什么

成功更新后,当安装有两个或更多 profile、至少一个次要 profile 跑自己的网关(活跃进程或已装服务)且 gateway.multiplex_profiles 关闭时,hermes update 运行同一预检:

  • 没有东西阻塞 → 迁移自动跑(与 hermes gateway migrate --multiplex --yes 同一代码路径)并打印它做了什么。这是确定性的、绝不提示,因此也在无头/cron 更新上跑。
  • 有东西阻塞 → 一个警告块列出每个阻塞及其确切修复和稍后要跑的一行命令。什么都不改。

单 profile 安装永不迁移(没收益),已多路复用的安装不动它。没有次要 profile 跑自己网关时 hermes update 也什么都不做——它绝不在什么都没跑的安装上翻模式。

hermes update 绝不自己跨过的边界

无人值守钩子只折叠共享一个 UNIX 用户、一个服务域和一个 profiles/ 树的 profile——hermes profile create 产出的形态。任何这些边界后的独立次要 profile 都会停掉自动路径:

边界示例
不同服务管理器或作用域默认在用户 systemd,次要在系统 systemd(或 launchd),或默认分离而次要由服务管理
一个 profile 装了多于一个 unit同一 profile 的用户 unit 和系统 unit(显式命令移除两者)
不同 UNIX 用户带自己 User= 的系统 unit,或另一个 uid 拥有的活跃网关;User= 这台宿主机无法解析的系统 unit——在次要或默认上——算未知,绝不算"同用户"
HERMES_HOME 在 <default home>/profiles/ 之外固定 HERMES_HOME=/opt/hermes/profiles/emma 的 unit

那种情况 hermes update 打印它发现的边界加 hermes gateway migrate --multiplex,什么都不改——不删 unit,按 profile 网关继续跑(边界挡住折叠处 --force 仍是路径;没有这些边界的 profile 用 gateway.standalone: true 退出)。折叠这样的机群是用进程内隔离替换内核强制边界(文件属主、User=),这是运维的决定。显式命令仍做得到:同样发现作为 notices 出现在 hermes gateway migrate --multiplex --dry-run,让你先读,确认后 --multiplex 继续。

退出自动迁移

在默认 profile 设 gateway.auto_multiplex_migration: false,让自动折叠永不在本安装跑:

hermes config set gateway.auto_multiplex_migration false

hermes update 随后让按 profile 网关保持原样,无输出无改动,无论安装看起来多合格。设置在配置里,因此跨更新存活——决定做一次,而非每次发行重新争论。它像其他设置一样从生效配置读取,因此固定在托管作用域(/etc/hermes/config.yaml)的值压过 profile 自己的文件。它只管自动路径:hermes gateway migrate --multiplex 是显式请求,仍会迁移(也是重新加入的受支持方式)。缺失或 true 保持上文默认行为。

显式命令不同:带两个或更多 profile、且没有独立次要网关时,hermes gateway migrate --multiplex 仍应用剩下那一步——设 gateway.multiplex_profiles: true 并(重)启动默认网关。你要多路复用;你得到多路复用。

克隆不带频道

hermes profile create --clone 留下源的 bot token 和允许列表(见 Profiles → 消息频道永不被克隆),因此克隆机群不再触发下面的重复凭据阻塞。仍带它们的旧克隆会被 hermes profile list 标记。

迁移做什么

  1. 停掉每个次要 profile 的独立网关并卸载其服务(systemd 用户/系统 unit 或 launchd agent)。移除了什么记录在 ~/.hermes/gateway_migration.json 以便回滚。
  2. 在默认 profile 的 config.yaml 设 gateway.multiplex_profiles: true。
  3. 重启默认网关——或在次要所用的同一服务管理器上安装并启动它,让 systemd 管理的机群保持 systemd 管理。
  4. 等默认网关记录覆盖每个 profile 的 served_profiles,然后打印摘要。

阻塞与修复

阻塞原因修复
两个 profile 配置了相同平台凭据(如同一 TELEGRAM_BOT_TOKEN)一个进程下 bot token 只能轮询一次;多路复用器会停驻重复项,该 profile 的 bot 会静默从第二个 profile 移除 token,或留在 default 里用 profile_routes 路由该 profile 的聊天
次要 profile 启用了一个在默认监听器上没有 /p/<profile>/ 入口的端口绑定平台多路复用器跳过整个 profile(见规则 2)在该 profile 禁用平台(platforms.<name>.enabled: false),或让 profile 独立运行:在其自己 config.yaml 设 gateway.standalone: true 并等宿主机重扫(最多 30 秒),或发其 rescan-profiles 控制动词。只在边界挡住折叠处用 hermes -p <name> gateway install --force。

凭据检查复用网关自己的冲突检测,因此其判定与多路复用器启动时一致。哪些端口绑定平台有 /p/<profile>/ 入口从适配器自己读取(每个声明 serves_profile_prefix),因此预检随新 HTTP 入站适配器获得前缀而保持正确。

入站端口 profile 变了什么

一个曾在自己端口上用 api_server 或 webhook 的次要 profile 不被阻塞——但其 URL 变了。预检打印确切新 URL,例如:

Profile 'coder': api_server moves onto the default listener at
http://127.0.0.1:8642/p/coder/v1/... (its key/secret is unchanged; update
clients that call the old per-profile port).

该 profile 自己的 API_SERVER_KEY / webhook secret 继续认证带前缀 URL;key 其他方面不变。

迁移后创建的 profile

多路复用器运行时创建的 profile 无需重启即被服务(见上)。活跃多路复用器捡起 profile 时 hermes profile create 确认这一点;只有够不到多路复用器时(例如从旧构建启动的网关)才打印 hermes gateway restart 提醒。

失败处理与续跑

迁移是事务性的。它能从计划预见的失败(一个必须以 root 运行却没有记录 User= 的系统 unit、一个它无法重写的配置文件)在停任何按 profile 网关之前拒绝。清单写入后任何失败——写标志、后来次要的停止或 unit 移除、默认的安装或启动——当场通过清单回滚,因此没有 profile 被留下没网关。万一进程在那个窗口任何地方死掉,下次 hermes gateway migrate --multiplex 看到标志开着、清单、且没有活跃多路复用器服务被迁移的 profile(一个已装但停的默认 unit 不算),就从清单续跑,而非报告"already multiplexed"。磁盘上的清单永远意味着未完成:它是续跑记录,不是回滚命令——没有 --standalone 反向,上面的补偿器只在单次失败应用内跑,因此没有 profile 被留下没网关。

不自动覆盖:s6 监管容器——它们在下次容器启动时收敛(按 profile 槽位注册为 down,根网关多路复用;gateway.standalone: true profile 从自己的运行意图启动自己的槽位)。Windows 计划任务由该命令折叠。仪表盘 System 页在预检发现合格安装时把同一迁移作为按钮提供。

更新代码

hermes update 拉一次最新代码并把新内置 skill 同步进每个 profile:

hermes update
hermes-gateways restart

运行中的网关由更新自己重启;在仍每 profile 跑一个网关的安装上,更新随后运行迁移到单多路复用网关——无阻塞时自动,否则作为警告点名边界(不同 UNIX 用户、HERMES_HOME 在 profiles/ 之外)和你自己跑的一行命令。

用户修改的 skill 绝不被覆盖。

故障排查

"Could not find service in domain for user gui: 501"

你在之前的 hermes gateway stop 后跑了 hermes gateway start。CLI 的 stop 做完整 launchctl unload,把服务从 launchd 注册表移除。CLI 在 start 时抓住这个特定错误并自动重新加载 plist(↻ launchd job was unloaded; reloading service definition)。服务正常启动。无需修复。

崩溃后陈旧 PID

如果一个 profile 的网关显示 not running 但进程还活着:

ps -ef | grep "hermes_cli.*-p <profile>"
cat ~/.hermes/profiles/<profile>/gateway.pid
kill -TERM <pid>          # 优雅
kill -KILL <pid>          # 几秒后还不行则
<profile> gateway start

强制硬重置一个服务

# macOS
launchctl unload ~/Library/LaunchAgents/ai.hermes.gateway-<profile>.plist
launchctl load   ~/Library/LaunchAgents/ai.hermes.gateway-<profile>.plist

# Linux
systemctl --user restart hermes-gateway-<profile>.service

健康检查

hermes doctor                  # 默认 profile
hermes -p <profile> doctor     # 一个 profile