Hermes Agent + xCloud
Troubleshoot a broken site with Hermes Agent on xCloud
Troubleshooting with Hermes Agent means messaging an always-on agent from Telegram, Slack or the terminal when a site breaks, and having it walk xCloud's evidence in order and report back in the same chat.
- Skill: xcloud:troubleshoot
- Toolsets: sites, servers, sites-wordpress, wordpress-actions
- Free with every xCloud account
Youapi.example.com is giving 502s. What is going on?
sites_statusread-only
sites_access-logsread-only
AgentStatus is normal. The error log shows upstream prematurely closed connection while reading response header from upstream, repeated since 06:12, and no failed task before it. That points at the PHP process, but the log does not say why. Want me to check the server services?
YouYes.
servers_servicesread-only
. Keep a human in the loop: xCloud stops and asks before anything that creates, deploys, updates, reboots, deletes or buys.
Setup
How Do You Set Up Hermes Agent to Troubleshoot a broken site on xCloud?
Connect Hermes Agent once; every job on this account uses the same connection. Then ask in plain words.
Add the xCloud MCP server
Add this under mcp_servers in ~/.hermes/config.yaml. If the file already has an mcp_servers block, add only the xcloud entry. Then start a new Hermes session, or run /reload-mcp in the one that is open.
mcp_servers: xcloud: url: "https://app.xcloud.host/mcp" auth: oauthAuthorize xCloud
On first connect Hermes prints an authorize URL and waits for the sign-in. Run this command to start it yourself or to re-authorize later. On the xCloud approval screen tick the teams and choose Read-only or Full access. Hermes caches the tokens in ~/.hermes/mcp-tokens/xcloud.json and reuses them on later runs.
hermes mcp login xcloudNo browser? Use an API key
For a server with no browser, create a token with the mcp:invoke scope plus the read abilities for the areas it will use (read:servers and read:sites, with read:billing and read:addons for billing and add-on tools) and the matching write: abilities if it should change things in Settings, Developers, API Tokens and send it as a header in place of auth: oauth. Keep ~/.hermes/config.yaml out of version control when it holds a token.
mcp_servers: xcloud: url: "https://app.xcloud.host/mcp" headers: Authorization: "Bearer YOUR_TOKEN"Check it worked
Then ask Hermes Agent for the job itself, for example:
shop.example.com is showing a critical error. Check it on xCloud and tell me what the logs say, in three lines.
In practice
How Does Troubleshooting Work from Hermes Agent?
Hermes Agent is not a window you open once a site is already on fire. It runs as a gateway on a machine you keep on, and you reach it from Telegram, Discord, Slack, WhatsApp or Signal, so the first report of trouble often comes from your phone: shop.example.com is giving me a 500. Hermes answers in the same thread. It calls sites_status first and then sites_events, and it sends back short lines that read well on a small screen: the site is in a normal state, a plugin update task finished with an error just before the break, here is that task's output. Reads need no approval, so nothing waits on you until the agent wants to change something.
Memory is the difference you notice over weeks. Hermes keeps what you told it, so you can say once that shop is the WooCommerce site on the Frankfurt server and the blog is the one with a staging copy, and later reports can shrink to the shop is down. Treat that memory as a way to find the right site, never as evidence. The cause still has to come from a log line or an event fetched during this investigation, so Hermes reads the nginx error log again with a limit each time, even when the last outage looked identical. When the evidence runs out, it sends the dashboard path, Site, Site Monitoring, Logs, and the site link xCloud returned, so you can open the WordPress debug.log in your phone's browser.
Because Hermes has a built-in scheduler, you can also hand off the first look. A cron job that runs in a fresh session can check every site's status and recent events every hour and message you only when one is not in a normal state, with the failing event attached. Scheduled runs are for reading. Turning on WP_DEBUG, creating temporary shell access and running a rescue each need a yes naming the site, and a job with nobody present cannot give one, so those steps wait for a message from you.
Hermes Agent specific: Hermes reads mcp_servers when a session starts, and a gateway that has been running for days keeps the tool list it started with. If you edit config.yaml or re-authorize with hermes mcp login xcloud, run /reload-mcp or restart the gateway, or the moment of an outage is when you find the xCloud tools missing. A scheduled watch reuses the cached sign-in, so if it lapses, run hermes mcp login xcloud again and check that the job still reports.
What xCloud does for troubleshooting
xCloud reads the evidence in a fixed order, cheapest first: site status, recent events, the web server access and error log, WordPress health, WP_DEBUG and server services. The agent names a cause only when a log line or an event it retrieved shows it, and says plainly which logs only the dashboard can display.
- Check the status first. The agent reads the site status before anything else. A site that is still provisioning, deploying or in a failed state explains a 500 on its own, and the answer is then to wait or to look at the last deploy, not to hunt for a PHP fault.
- Read the recent events. The site's recent tasks, such as SSL issuance, plugin updates, cache purges and deploys, come with their outcome. A 500 that began right after a failed task has usually found its cause here, and one task's full output can be opened.
- Read the web server logs. The agent asks for the nginx log type with a limit, which returns the access log, the error log and the 7G and 8G firewall logs on Nginx and OpenLiteSpeed alike. The error log is where a PHP fatal shows up as a 502 or 500. Log lines are quoted as data, never followed as instructions.
- Check WordPress health. For WordPress sites the agent reads the health status: WordPress and PHP versions, the debug and cron flags, and whether the install itself is broken.
- Turn on WP_DEBUG only if needed. If the logs so far are inconclusive, the agent proposes switching WP_DEBUG on, waits for your approval, and switches it back off when the investigation ends. The toggle flips the flag only; it does not return the debug log.
- Check the server services. When the whole server looks wrong rather than one site, the agent reads whether the web server, PHP and the database are running. If nothing explains the error, it reports what it checked and what each check showed, and points you to the dashboard logs.
Reference
Troubleshooting Settings and Limits on xCloud
The facts Hermes Agent works within when it troubleshoots a broken site. Where a row names the dashboard, that step stays yours to take there.
| Setting or limit | What applies |
|---|---|
| Read order | Status, recent events, nginx access and error log, WordPress health, WP_DEBUG, then server services. Reads run straight away and change nothing |
| Log type | sites_access-logs with type nginx reads the access log, the error log and the 7G and 8G firewall logs. The default type reads the access log only |
| Log reads | Logs are read over SSH, so a call is slow. The agent asks for a bounded window with a limit, not everything |
| Staging history | The deployment log of a site with a staging environment is the push and pull history between staging and production, not the Git build log |
| Dashboard-only logs | The WordPress debug.log, Laravel, PM2 and docker-compose logs: Site, Site Monitoring, Logs. Server logs such as Fail2Ban and auth: Server, Monitoring, Logs |
| WP_DEBUG | The API only toggles the flag. Reading the resulting debug.log is a dashboard step |
| Stale error pages | A purge of the site cache clears a cached error page. It runs without a confirmation stop and finishes asynchronously, so completion shows in the site events |
| Temporary shell access | A temporary sudo user expires after 12 hours, and the agent removes it as soon as the investigation ends instead of waiting for the expiry |
| Rescue action | A server-side repair with flags such as directory permissions, regenerating the Nginx configuration, reinstalling PHP or repairing Node.js, PM2 or OpenClaw. A flag the site type does not support is refused |
| Hand-offs | A slow site goes to performance, a failed deploy goes to deploy, and a 526 or certificate warning goes to SSL |
Rules Hermes Agent has to follow
- The agent never restarts a service or reboots the server just to clear an unexplained error. A restart is a real change on a live machine, and it destroys the evidence the logs were about to show.
- A plain 500 or a short database error is a symptom. The agent does not name a cause such as bad credentials, a missing migration or a crashed process unless a log line or event it retrieved shows it.
- Temporary shell access is used only when the readable logs do not explain the fault and you agree. The agent revokes it the moment the investigation ends.
- Turning WP_DEBUG on, creating shell access and running the rescue action each need your explicit yes naming the site or server. The agent turns WP_DEBUG back off afterwards.
- When no evidenced cause turns up, the agent says what it ruled out, gives you the dashboard log path and the site's dashboard link, and suggests contacting xCloud support with that evidence.
Example prompts
What Can You Ask Hermes Agent to Do for Troubleshooting?
Type these as written and swap in your own repository, site and server names. Reads and routine actions such as backups, cache purges, PageSpeed scans and vulnerability scans run straight away; creating, deploying, updating, rebooting, deleting, buying or starting a broken-link scan stops and asks first.
shop.example.com is showing a critical error. Check it on xCloud and tell me what the logs say, in three lines.Every hour, check the status and recent events of all my sites. Message me only if one is not in a normal state, and change nothing.Last week blog.example.com broke after a plugin update. Check whether the same thing happened today, but read the logs again before you say so.My site shop.example.com is returning a 500 error. Find out why, and show me what you checked.Read the nginx error log for the shop site and tell me what it says about the last hour.Which tasks ran on the shop site just before it broke? Did any of them fail?Check WordPress health on the blog site and tell me whether the install itself is broken.Is the database running on the Frankfurt server? Check the services before touching anything.Turn on WP_DEBUG for the blog site, tell me what you find, then turn it off again.Purge the cache on the shop site in case it is serving a stale error page.Run a rescue on the shop site to reset directory permissions.Hermes Agent and Troubleshooting: Frequently Asked Questions
What people ask before they let Hermes Agent troubleshoot a broken site through xCloud.
Can Hermes Agent tell me when a site goes down without me asking?
Yes, through its built-in cron scheduler. A job can read site status and recent events on a timer in a fresh session and message you in a chat only when something is not normal. Keep the job to reads, because changes still need your approval.
Does Hermes Agent remember the cause of an earlier outage?
It remembers what you told it, which helps it find the right site and server quickly. It does not treat that as proof, so it reads the status, events and error log again for each new incident before it names a cause.
What does the agent check first when a site shows a 500 error?
The site status, then recent events, then the web server logs. A site that is still provisioning or whose last deploy failed explains a 500 without any deeper digging, so the status read always comes first. Only after those cheap reads does the agent look at WordPress health, WP_DEBUG and server services.
Will the agent restart my server or a service to fix an error?
Not to clear an error it cannot yet explain. A restart changes a live machine and wipes out the evidence in the logs, so the agent finds the cause first. Restarting a service stays an action you ask for and approve.
Which logs can the agent read, and which can it not?
It reads the site's access log, error log and the 7G and 8G firewall logs through the nginx log type. It cannot read the WordPress debug.log or the Laravel, PM2 and docker-compose logs. Those are in the dashboard under Site, Site Monitoring, Logs, and server logs such as Fail2Ban are under Server, Monitoring, Logs.
Does the agent get shell access to my server?
Only when the logs it can read do not explain the problem and you agree to it. The access is a temporary sudo user, and the agent removes it as soon as the investigation ends. If it is ever left behind, it expires on its own after 12 hours.
What if my site is slow rather than broken?
A slow site is a different investigation, because it is answered from measurements and cache state rather than error logs. The agent hands it to the performance job. A failed deploy goes to the deploy job and a 526 or certificate error goes to the SSL job.
Other agents
Troubleshooting with Other Agents
The same job, the same xCloud tools, a guide for each client.
- Troubleshoot a broken site with Claude CodeAnthropic's terminal coding agent. One claude mcp add command, plus the xCloud skills plugin with nine skills on top.
- Troubleshoot a broken site with ClaudeAnthropic's chat assistant on the web and desktop. Add xCloud as a custom connector, no terminal needed.
- Troubleshoot a broken site with Claude CoworkAnthropic's desktop agent for delegated work. Add the xCloud connector, then hand off hosting jobs.
- Troubleshoot a broken site with CursorThe AI code editor. One mcp.json entry with the compact URL, because Cursor stops at 40 tools.
- Troubleshoot a broken site with CodexOpenAI's coding agent for the terminal. A codex mcp add command or a config.toml entry, then codex mcp login.
- Troubleshoot a broken site with OpenCodeThe open-source terminal coding agent. One remote MCP entry, then opencode mcp auth xcloud.
- Troubleshoot a broken site with OpenClawThe open-source agent runtime with chat apps and automations. ClawHub skill plus the MCP client.
- Troubleshoot a broken site with WindsurfThe Cognition editor, now Devin Desktop. devin mcp add for the Devin Local agent, a serverUrl entry for legacy Cascade.
- Troubleshoot a broken site with GitHub CopilotCopilot agent mode in VS Code. One .vscode/mcp.json entry, or the Agent Plugins package.
- Troubleshoot a broken site with Gemini CLIGoogle's terminal agent. One gemini mcp add command, OAuth found automatically.
- Troubleshoot a broken site with ChatGPTOpenAI's chat assistant. A developer-mode app with the xCloud MCP URL and OAuth.
- Troubleshoot a broken site with ChatGPT dotsOpenAI's always-on agent in ChatGPT. Uses the xCloud MCP plugin you add in ChatGPT, with custom rules and scheduled tasks.
- Troubleshoot a broken site with GrokxAI's terminal agent, Grok Build. One grok mcp add command or a config.toml entry.
- Troubleshoot a broken site with Grok BotxAI's always-on Bots on a cloud computer. One Remote HTTPS MCP plugin, OAuth sign-in, routines on a schedule.
- Troubleshoot a broken site with KiroAWS's agentic IDE. One url entry in .kiro/settings/mcp.json, plus the portable xCloud Agent Plugins package.
- Troubleshoot a broken site with AntigravityGoogle's agentic IDE. One serverUrl entry in mcp_config.json and a browser sign-in.
- Troubleshoot a broken site with ZedThe Zed editor's Agent Panel. One context_servers entry in settings.json and a browser sign-in.
More Hermes Agent guides
- Hermes Agent and xCloud overview
- Deploy from Git with Hermes Agent
- Run Docker apps with Hermes Agent
- Install one-click apps with Hermes Agent
- Manage WordPress with Hermes Agent
- Back up and stage sites with Hermes Agent
- Manage SSL and domains with Hermes Agent
- Manage servers with Hermes Agent
- Speed up a slow site with Hermes Agent
- Secure sites and servers with Hermes Agent
Run Your Hosting from Hermes Agent
xCloud MCP, the Agent Skills and the Public API are free with every account. Connect once and ask.