Separate rush-client and rushx-client binaries, opt-in until cutover. Existing
rush, rushx, and their reporter entrypoints are unchanged.
To try the daemon on the rushstack repository itself before a release contains it, follow the contributor dogfooding guide.
The dependency on @microsoft/rush is intentional: it is the version-selecting
frontend, whereas @microsoft/rush-lib is the execution engine. Native routing
and safe daemon fallback in src/launchClient.ts restore the rush or rushx
executable name and load the frontend's exported lib/start entrypoint. They
must not call the client's bundled Rush.launch() or Rush.launchRushX()
directly: the workspace or RUSH_PREVIEW_VERSION can select a different engine.
The client also reuses the frontend's MinimalRushConfiguration for native
Rushx discovery output instead of maintaining a second implementation.
The startup sequence in apps/rush/src/start.ts and RushFrontend.ts selects
and initializes the reporter before RushVersionSelector installs or loads
the selected engine. RushCommandSelector then handles engine capabilities
and old-engine output compatibility. Selection supports releases predating
rush-lib (before Rush 4), and cannot depend on a new API being present in the
selected engine. Calling a selector alone would also bypass reporter startup,
preview-version validation, and native launch options.
The runtime dependency direction is client → frontend → engine; rush-lib
does not depend on this client or on the frontend. Re-exporting the existing
frontend selector from rush-lib would introduce an engine → frontend → engine
cycle. A future standalone bootstrap package would need to own the complete
startup/reporting contract, not just engine installation. Until that separation
is warranted, reusing the existing frontend preserves version selection and
reporter behavior without duplicating or relocating bootstrap code.
Routing precedence:
--no-daemonbefore--, help, never-daemonize commands, an option before the command other than--quiet/-q(such as--debug), and Rushx with any TTY stdio stay in-process.--quietand-qonly hide native Rush's startup banner, which a daemon request never prints, sorush -q buildis sent to the daemon without them.- CI stays in-process unless
RUSH_DAEMON=1explicitly opts in, even if config enables the daemon. RUSH_DAEMONoverridesrush.json'sdaemon.enabled; the default is false.- Auto-start is considered only after selecting daemon execution.
A command that routing keeps in-process shows no progress line. It says why in one stderr line before native
Rush starts, in agent mode (see Output modes) and, in legacy mode, when RUSH_DAEMON=1 asked
for the daemon. Rushx uses the rushx-client: prefix. --no-daemon and help print nothing. The lines are:
rush-client: RUSH_LOG_LEVEL selects the native reporter; using in-process Rush.The same line namesRUSH_REPORTER=<value>,--reporter,--output,--log-leveloruseRushReporter in experiments.json(see below).rush-client: the daemon does not support "--debug"; using in-process Rush.rush-client: the daemon does not run "check"; using in-process Rush.rushx-client: the daemon does not run scripts in a terminal; using in-process Rush.rush-client: RUSH_DAEMON=0 turns the daemon off; using in-process Rush.rush-client: the daemon is not enabled for this repo; using in-process Rush.To enable it, setdaemon.enabledtotrueinrush.json, or setRUSH_DAEMON=1in the environment.rush-client: CI is set, so the daemon is off unless RUSH_DAEMON=1; using in-process Rush.The line names the first CI marker that is set:CI,TF_BUILD,GITHUB_ACTIONS,JENKINS_URLorTEAMCITY_VERSION.
When the selected daemon cannot be reached or started, an ordinary invocation prints the reason and runs
in-process (rush-client: <reason>; using in-process Rush.). A startup failure while a live process can
still make the daemon ready is the exception: a process that listens at the endpoint but does not
complete hello/ping in time, a startup helper that still waits for its daemon, or another client that
holds the start mutex, as in a burst of clients that all find no daemon. In-process Rush would take the
repository lock, and the requests that the daemon serves would then wait for it, or fail when their wait
timeout ends (see below). Instead, the client keeps trying for one more startup deadline
(15 seconds, so about 30 seconds in all) and uses the daemon once it is ready. It says so when it starts
waiting (rush-client: The daemon is not ready yet. <live process>, so this command waits up to 15 s more for it instead of running Rush in-process.; agent output shows "rushd is still starting; waiting for it"
as the progress phase, and on a pipe writes it as a progress line that ends with because <live process>). If the daemon is still not ready, the command exits with code 1. The message
gives the startup error with its --no-daemon hint, then the process that is still live, "so Rush was
not run in-process", and a pointer to rush-client daemon status. When the process that the daemon's
ownership record names still runs but does not answer (on Linux, for example because a signal stopped
it), the message instead says what that process is doing, then on a line of its own that Rush was not
run in-process, and its last line says what to do, for example Resume it with "kill -CONT <pid>"; it then serves the next command. On Linux, when that process has this workspace's ownership record open, as the daemon that
wrote it does, and stays stopped (state T or t) while the client samples it for 1.5 s, the command does
not wait for either deadline: it exits with code 1 and that message once the 1.5 s have passed. It names
a signal to send only to a Rush daemon that has that record open; for any other process it says to end
that process if it is this workspace's daemon, and else to delete the ownership record, and that until
then each command that uses the daemon first waits 15 s for a daemon to answer. Such a Rush
daemon that still runs after its socket file was deleted also fails the command this way, with code 1
and without running Rush in-process, once the client has waited 15 seconds for it to exit; its last line
says that it may exit once its running requests finish. When the process that the ownership record names
has exited but is not reaped yet (on Linux, state Z), nothing live can make the daemon ready, so the
command runs Rush in-process once its startup deadline (15 s) has passed. The last line of its message
names the parent that has not reaped that process, and says that until then each command that uses the
daemon first waits 15 s for a daemon to answer.
When the client runs Rush in-process after it tried the daemon, because it could not reach one or because the
daemon handed the request back, and another Rush process holds the repository's lock, such as the daemon while it
builds for another request, Rush waits for the lock instead of failing at once with "Another Rush command is
already running in this repository." It waits only for what is left of the request's wait timeout (see below),
counted from when the client sent the request, or from when it gave up on the daemon if it could not reach one;
the built-in 30-second default applies. It writes one stderr line when it starts to wait (Waiting up to 28 s for the Rush daemon (PID 4242) to release this repository's lock.). If the lock is still held at the deadline, the
command fails as before, and the error names the holder (The Rush daemon (PID 4242) still holds this repository's lock.). With --no-wait or a zero timeout, Rush tries once and names the holder. On Windows the lock file does not
name a process, so Rush says "another Rush process". Rush that routing keeps in-process (such as --no-daemon,
RUSH_DAEMON=0 or CI), rushx-client, and a workspace that selects another Rush release than the one the client
bundles fail at once, as native Rush does.
--no-wait fails immediately when daemon admission is unavailable.
--wait-timeout SECONDS (or --wait-timeout=SECONDS) overrides the configured queue
timeout; finite nonnegative decimal seconds up to 2147483.647 are accepted and
rounded down to milliseconds. These controls are mutually exclusive and are
consumed before forwarding, never appended to a project script. Arguments after
-- remain literal script arguments.
Only time that the request spends waiting for other requests counts against the
queue timeout, whether it comes from --wait-timeout,
RUSH_DAEMON_QUEUE_TIMEOUT_SECONDS, daemon.queueTimeoutSeconds in rush.json,
or the built-in 30-second default. Waiting while another request
loads or reloads the workspace graph does not count, so every build that arrives
while the first build after startup loads the graph runs once the load finishes.
That wait fails after 10 times the timeout (5 minutes with the default), so a load
that never finishes does not hold other requests forever. The request's own work,
such as checking its inputs, loading the graph, routing and execution, does not
count either. A configured or per-invocation timeout also
limits waiting for a running build to start so that the request can join it (with
joinRunningBatch), waiting for a running build that the request could not join, and waiting for
the requests that the daemon is serving to finish before it restarts for the
request's environment. The built-in default does not: with it, a build that arrives
while a compatible build is already running waits for it to finish and then runs,
instead of failing after 30 seconds, and a request that needs a restart waits for
the requests that were running when it arrived to finish and then runs on the
restarted daemon. The default still limits a restart wait while the daemon runs a
rushx script, such as a dev server, which may not exit until it is stopped, and
while it serves requests that arrived later. A rushx-client script that arrives
while another request waits for the daemon to restart does not start on the old
daemon, where the restart would wait for it to exit: it waits for the restart and
then runs on the restarted daemon, and its timeout applies to that wait as it does
to the restart wait. --no-wait fails wherever the request would wait. On a
timeout, the client exits with code 1 and suggests --wait-timeout. It does not
suggest exporting RUSH_DAEMON_QUEUE_TIMEOUT_SECONDS, because Rush versions that do
not recognize a RUSH_ environment variable fail every command while it is set.
In legacy output and in rushx-client, the admission failure line
(rush-client: daemon admission failed (wait-timeout): …, or (no-wait); in
rushx-client it begins with rushx-client:) gives the daemon's reason, as agent
mode's summary line does, so it names what the request waited for, such as a daemon
restart, and why the daemon restarts.
A request also waits while a Rush process that the daemon does not run, such as
rush install or a --no-daemon build, holds the repository's lock. Every timeout,
the built-in default included, limits that wait, since that process can run for any
length of time. Stderr, on a terminal and on a pipe, names the process at once
(rush-client: waiting for another Rush process (PID 12345: rush install) to release this repository's lock.), and again with the time waited every 10 seconds
(still waiting after 10s for …); agent output shows it as the progress phase, and if the
daemon then restarts, request resubmitted to the new daemon; preparing the workspace graph
until the new daemon reports a queue position or starts the command. Only
Linux tells the PID and command; elsewhere the line says another Rush process. On
Linux, the name also says when that process is stopped, since it cannot release the
lock until something resumes it (another Rush process (PID 12345: rush install; it is stopped (state T), for example by SIGSTOP)).
--no-wait and --wait-timeout 0 fail at once, and the admission failure line of
these and of a timeout names the process. A request never waits for a lock that the
daemon itself holds for another request: it fails at once, as before.
Admission controls also apply to experimental graph requests, but not
start|stop|restart|status|logs. They affect daemon admission, and how long Rush that runs in-process after
the client tried the daemon waits for the repository's lock (see above); otherwise native fallback
retains native command behavior. Waiting positions are shown on interactive stderr,
and a wait for a daemon restart (see below) or for another Rush process on a pipe
too. Admission failures report their typed reason and a nonzero exit code.
Explicit reporter/output/log-level controls (--reporter, --output, --log-level,
RUSH_REPORTER other than legacy, or RUSH_LOG_LEVEL) retain the native frontend
reporter path, with or without --no-daemon, including --reporter=ai. The daemon client
does not silently reinterpret requests for JSON, AI, file, or other reporter formats.
A repository that opts into the native reporter with "useRushReporter": true in
common/config/rush/experiments.json also stays on the native (in-process) path, so
that its reporter output is honored rather than silently replaced by the daemon
stream. Native reporter rendering over the daemon protocol is a follow-up.
The rush-client daemon path has two output modes (rushx-client always uses legacy).
Requests that use the native reporter path (see above) always get native output, and
agent mode writes nothing ahead of it. Otherwise, selection precedence is:
RUSHD_OUTPUT=agentorRUSHD_OUTPUT=legacy.- An active
COPILOT_CLIagent marker selectsagent, matchingdetectAgent()in@rushstack/reporter(a value is inactive when empty,0,false,nooroff). Other agents can opt in withRUSHD_OUTPUT=agent. - Otherwise
legacy: the unchanged collated operation stream.
Agent mode is plain text for humans and agents, not the AI reporter's JSON record format;
use --reporter=ai for machine-parsed records. On a TTY it paints at most three live rows, the
first before @microsoft/rush-lib is loaded. On a pipe it writes one progress line when the
daemon has the request (rush build · 0.1s · sent to rushd; preparing the workspace graph (status at least every 25s)), however many operations run. A longer request also gets status
lines, so that it does not look hung: one whenever nothing was written for 25 s, with the counts
and the running operations, and one when connecting to the daemon takes more than 10 s. A
wait for a daemon that is still starting gets a line of its own
(rushd is still starting; waiting for it (up to 15s more) because its startup helper (PID 4242) is still waiting for the daemon). A
request that waited for admission says so at the end of its summary line
(· queued behind another request (position 1 at 0.2s)), unless it failed: a failure's summary
line gives the failure, not the wait. It always ends with one summary
line, for example
rush build: SUCCESS 772/772 operations (12 success, 760 from cache) in 3.1s, or
up to date (no operations needed), or, when the selection parameters matched no projects,
rush build: SUCCESS 0 operations in 0.5s · the selection parameters did not match any projects.
The counts follow the native summary: silent operations
(such as phases a project does not define) are not counted unless they fail, and operations
that did not need to run (SKIPPED or NO OP in native output) are counted as up to date. The verdict is
SUCCESS, FAILURE or CANCELLED (Ctrl+C or a termination signal); a request that a daemon
shutdown aborted is a FAILURE whose summary line gives the reason, with exit code 1. When the
daemon did not admit the request in time, or at once with --no-wait, the reason on the summary
line starts with daemon admission failed (wait-timeout) or daemon admission failed (no-wait),
as in legacy output, followed by the daemon's reason in full. Warnings and errors that Rush or a
Rush plugin writes outside any operation (for example a plugin that continues without the cloud
build cache) are written at the end, at most three lines of them, before the summary line and any
operations reported with it.
A failed operation is reported as soon as it fails, while the rest of the request runs on: a
failed: <operation> · full log: <path> line and a short excerpt of its output, error lines
with the line that follows them first, then the last lines. Stack frames, Require stack: lists
and progress noise are left out, and so are a message that a tool repeats in its summary, an
error count that the shown errors account for, and, when the first error shown names a source
location, the lines before it. Up to three operations are reported. Two kinds are reported just
before the summary line instead: a failed operation that wrote no output, with the error from the
daemon's result, and, when no operation failed, the operations whose warnings failed the request
(warnings: …). An operation that succeeded and then got warnings, because its build cache entry
could not be written, is shown with the output that it wrote after it succeeded. The error of a
reported operation that wrote output is printed too, unless its excerpt shows it or it only gives
the exit code (Returned error code: 1): for example an error
thrown while the operation's build cache entry was restored. Only the daemon's result carries it,
so for an operation reported as it failed it comes just before the summary line, as an
error: <operation> line followed by the error. On a pipe, a status line names a failed
operation that wrote no output 1 s after it failed, unless the result came first. The summary
line names up to five failed (or warning) operations. Every operation's full output is in its
project's rush-logs/ folder, whether or not it was printed.
When the daemon ran an operation's incremental command (its <phase>:incremental script; see
incrementalBuilds below) and that command failed, the line reads
failed: <operation> · incremental command; its next run uses the initial command · full log: <path>.
That command can fail where the initial command, which --no-daemon runs, would not. Warnings
that an incremental command reported get · incremental command in the same place.
When a request falls back to in-process Rush, agent mode stops and native output follows.
In agent mode a failed rush build doesn't wait for all of its work. Its result comes once an
operation failed and none of the selected projects that no other selected project depends on (for
example, the projects named by --to) is still waiting or running. The daemon keeps running the
operations that the failure didn't block, so that the next build finds them done, and the summary
line counts them and names up to three, in name order
(· 2 independent operations continue in rushd: lib-b (build), lib-c (build)). A later rush build
waits for them. While it waits only for them, its output says so and names up to three of them.
In agent mode the phase reads
queued behind 2 operations left running by an earlier failed command (position 1): lib-b (build), lib-c (build),
a status line on a pipe reads
waiting for 2 operations left running by an earlier failed command (queue position 1 at 0.1s): lib-b (build), lib-c (build),
and the summary line ends with
· queued behind 2 operations left running by an earlier failed command (position 1 at 0.1s): lib-b (build), lib-c (build).
Legacy output on a terminal prints
rush-client: waiting for daemon admission (position 1) behind 2 operations left running by an earlier failed command: lib-b (build), lib-c (build).
The daemon reports the position again each time one of them ends, so the phase and the status
line name only the ones that still run; the summary line keeps the first position that named them.
rush rebuild, rush install and rush update, a restart of the daemon for another
environment, and rush-client daemon stop stop them instead. So does a command that the daemon
doesn't run (such as a custom command that it can't serve), before the client runs it in-process;
the client then prints, indented under its fallback line,
rushd stopped 2 operations left running by an earlier failed command (lib-b (build), lib-c (build)), so that this command can run in-process.
A served command that makes the daemon reload its graph, such as rush test after rush build,
stops them as well, and so does a served phased command that isn't incremental, such as a custom
rush retest. These commands and a served rush rebuild name them while the daemon stops them:
the agent phase reads
stopping 2 operations left running by an earlier failed command (position 1): lib-b (build), lib-c (build),
a status line on a pipe reads
waiting while rushd stops 2 operations left running by an earlier failed command (queue position 1 at 0.1s): lib-b (build), lib-c (build),
and if the command succeeds or is cancelled, its summary line ends with
· stopped 2 operations left running by an earlier failed command (position 1 at 0.1s): lib-b (build), lib-c (build).
Legacy output on a terminal prints
rush-client: waiting for daemon admission (position 1) while rushd stops 2 operations left running by an earlier failed command: lib-b (build), lib-c (build).
Older clients say that such a command is queued behind them. Older daemons don't name them, so the
summary line ends with · queued behind another request (position 1 at 0.1s).
Rushx scripts, and built-in commands that only read the workspace (such as rush list), run
in-process alongside them, like two Rush commands at once in one checkout.
Older daemons report the failure when all of the work has ended.
Positively identified built-in install and update follow the same opt-in routing
precedence as workspace builds and require protocol 0.10
(DAEMON_WORKSPACE_RESTART_PROTOCOL_MINOR). They are not submitted to older peers.
Other package mutation, publishing, setup, management, and administrative commands
remain native rather than being forwarded as execution requests. Daemon management
subcommands use their separate control path. Rushx script names are not interpreted
as Rush built-ins. Arguments after -- are preserved.
Request cwd, environment, argv, width and color are captured before connecting.
The protocol currently expresses request color as a boolean; subscriptions carry
the corresponding color level. There is no SIGWINCH forwarding.
The standalone host now binds native build/rebuild requests, and the phased
commands of command-line.json (such as test), to a reusable all-project graph.
Global commands still run in-process. Native Rush parsing, project selection, graph plugins, and
incremental/cache semantics are reused rather than spawning another Rush CLI.
The client renders operation headers, collated text, and activity events; global
command byte streams remain byte-preserving. A rushx build script never claims
to be a workspace build.
Rushx requests carry invocationKind: "rushx" (protocol 0.8), independently of custom
command origin. Native parsing recognizes -q, -d, and --ignore-hooks before the
command; subsequent flags and -- belong to the script, apart from this client's
explicit admission/escape controls. Older peers fall back before receiving the request
or consuming input.
The standalone Rushx client conservatively keeps every script with TTY stdin, stdout, or stderr on the native path, before connecting, auto-starting, or consuming input. Arbitrary scripts may require raw mode or a controlling terminal, and a pipe is not a PTY. There is no attempt to discover this requirement by running and retrying a script. Fully non-TTY invocations retain daemon forwarding, binary byte parity, and EOF handling. An embedded client that knows its script is pipe-safe can still submit a Rushx envelope directly; it must declare any controlling-terminal requirement.
The default daemon installs RushDaemonRequestResolver(existingRushResolver) to enable real
package-script execution alongside native workspace builds. It reuses native Rushx parsing, escaping,
banner/diagnostics, lifecycle PATH and INIT_CWD preparation, dotenv precedence, and
pnpm injected-dependency synchronization. Only the actual script shell is spawned;
there is no Rush CLI child or synthetic warm graph. Native configuration discovery
is captured by the client and emitted only with daemon output, avoiding duplicate
discovery messages on fallback.
On Windows, Rushx retains native invocation-path spelling in script cwd, lifecycle variables and
pnpm-sync output, including 8.3 aliases and junctions. Daemon identity and confinement remain physical;
an alias does not create another workspace identity or permit execution outside the workspace.
Alias retargeting while queued fails before execution. Relative RUSH_TEMP_FOLDER initialization
uses safe in-process fallback.
The composite is exported for embedded hosts and wired into the standalone daemon.
No default or cutover flag is flipped. Active Rushx hooks, encrypted dotenv
vaults, changed process-global Rush configuration variables, stale workspace configuration,
and native help require pre-execution fallback. Ignored/recursive hooks retain native
behavior, including skipping post hooks after failure. PTY requirements remain in-process.
For forwarded Rushx, TTY output color/width overrides are applied after native
lifecycle environment preparation using the existing terminal policy helper.
Non-TTY requests retain their explicit environment values, including intentional
FORCE_COLOR/COLUMNS settings; no ambient daemon environment is merged in.
Compatible requests reuse the same native graph. Source changes refresh inputs;
changed configuration or command shape replaces the session and graph in the same
process. Environment, installed dependencies, implementation content, or selected
Rush version changes require a process restart rather than patching the existing
engine. Direct, inherited, and rig-based project configuration uses private native
loaders and is rechecked before execution. External plugins that participate in the
requested command (unassociated plugins, plugins associated with it, or plugin command-line
files that define it, its phases or parameters for either), .env, phased
watch/install options, and unsupported event-hook scripts still use typed
pre-execution fallback; plugins scoped only to other commands are permitted. This does not exclude the built-in install and update
commands described above. The native Rush lock is held for preparation and each
coalesced iteration, not while idle; native commands and --no-daemon can run
after a completed request without stopping the daemon.
Native workspace dispatch copies the request envelope and normalizes only the
engine-owned _RUSH_LIB_PATH to this daemon's own engine, keeping the spelling
that the engine chose when it loaded. Foreign client SDK paths therefore neither
select the wrong SDK nor cause a false restart. All other environment inputs
remain unchanged and participate in normal lifecycle checks.
Protocol 0.10 permits a bounded retry only when a pre-execution command result
explicitly carries retryAfterRestart: true. executeWithDaemonRestartAsync
waits for old ownership release and a validated successor, then resubmits an eligible
request at most once. Command input/output or cancellation prevents retry,
even with the typed flag. Unknown rejections and connection loss never authorize replay.
A started install or update is never repeated, including after a nonzero exit;
only an unstarted request can receive
the typed retry authorization. Accepted queued requests drain their typed restart
results before the old connection closes.
When the daemon's own installation was removed or replaced (for example a deleted
snapshot folder or a reinstalled Rush release), the daemon lets its running requests
finish, answers each other request with that typed restart once they have, and then
exits. The timeout rules of a restart for the request's environment apply (see
above): the built-in default does not limit waiting for the requests that were running
when the command arrived, but still limits it while the daemon runs a rushx script,
and --no-wait and an explicit --wait-timeout limit the whole wait. A command that
times out exits with code 1, names the changed folder and, if a script runs, suggests
stopping it. Otherwise the client starts a daemon from its own launcher once the wait
ends, resubmits the request, and prints one line on stderr (or above the agent
progress rows): rush-client: The daemon's installation at <folder> was removed; restarted the daemon (PID <pid>).
While a command waits for a daemon restart, for its installation, the command's
environment or the workspace's inputs, the agent progress status (or stderr, on a
terminal and on a pipe) says what it waits for and why the daemon restarts, as soon as
the daemon reports the wait: rush-client: waiting for 2 running requests to finish, including 1 rushx script; the daemon (PID <pid>) then restarts, because common/config/rush/pnpm-lock.yaml changed. After because, the cause is its installation at <folder> was removed (or replaced), this request's environment differs from the daemon's in NODE_OPTIONS (variable names, never their values),
<files> changed for the workspace's installation, the code of Rush or a Rush plugin changed (<files>), or this request selects Rush <version>. The count names rushx
scripts, because a script such as a dev server may run until it is stopped. A
rushx-client script that waits for another request's restart prints rushx-client: waiting for the daemon (PID <pid>) to restart for another request (2 requests ahead), because <cause>. The line goes to the stderr that the script writes to, which is a
pipe, because a rushx-client with a terminal runs the script in-process. Its first line
comes once rushx-client itself has started and sent the request, which takes about half a
second on a busy machine. A native install or update restarts the daemon once it ends,
unless it fails before it changes the installation (for example on the Rush lock), and
that restart would end the rushx scripts that the daemon runs. The daemon can't tell
beforehand whether the command will fail, so it first waits for them and says so the
same way: rush-client: waiting for 1 running rushx script to finish, since this command restarts the daemon (PID <pid>), which would end it. A terminal
gets a line whenever the wait changes, and a pipe when the wait begins or its cause
changes. Both get the line again with the time waited (still waiting after 25s for …) whenever 25 seconds pass without one, until the command follows the restart,
starts, or ends. Once the client asks rushd to cancel the command (Ctrl+C) and says so,
it writes no more wait lines. In agent mode, the progress phase says it, and on a pipe
a status line is written at once when the wait begins or its cause changes.
When the daemon restarts for a command's environment, the client prints a line of the
same kind that names the variables that differed, never their values:
rush-client: A command's environment differed from the daemon's in NODE_OPTIONS; restarted the daemon (PID <pid>). It names at most four variables and then says how
many more differed (A, B, C, D and 2 more). A command that was waiting when another
command's environment restarted the daemon prints that command's variables, because the
successor starts with that command's environment. A daemon that does not name the
variables gets no line. A variable name that holds a control character, such as a newline
or ESC, is printed with that character written as a \xHH escape, so the line stays one
line.
If the daemon that replaces it does not start, after either kind of restart, the command
fails with exit code 1 and one line that gives the reason for the restart before the
startup error: rush-client: A command's environment differed from the daemon's in NODE_OPTIONS; the restarted daemon did not start: <startup error>. A variable that keeps
the daemon from starting is then among the names in that line.
The client leaves out the …; restarted the daemon (PID <pid>). line of either kind when
the command already wrote a wait line (see above) for the same cause since it last
restarted, because that line said why the daemon restarts: the same variables, or the same
change to the same installation. The wait line counts when it was written as a line, on
stderr or as an agent status line on a pipe. With agent output on a terminal, the wait is
only in the live rows, so the line is printed above them.
When the connection is lost before a command's result, the command fails with exit code 1
and is not retried. The diagnostic keeps "Daemon disconnected before delivering a result; the
command was not retried." and says what happened to rushd. If its process exited (a crash, an
out-of-memory kill or a signal), it names the PID, points to rush-client daemon logs and, if
the daemon exits again, to --no-daemon (rushx-client --no-daemon for Rushx), and quotes on a
second line the fatal error that the launcher log recorded after the command was sent. Before it
prints that, the client removes the exited daemon's ownership record and socket, as the next daemon
start would. On Linux, it first stops the operations that the daemon left running, so a rerun, with
or without --no-daemon, does not race them.
Rush run in-process, with --no-daemon or as a fallback, first does the same when the ownership
record names a daemon that no longer runs, for example when the client that ran the command was
killed along with the daemon. On Linux that includes a daemon that has exited but is not reaped yet:
the client waits up to 1 second for it to be reaped, and else runs Rush without the reclaim, which a
later command does once the daemon's parent reaps it. Each of these reclaims, and the ones that daemon start, an automatic
start and daemon stop --force do, prints one line that says what it stopped, for example:
rush-client: Stopped the operations that the exited daemon (PID 4242) left running (process group 4242).
The line begins "Killed" instead, and ends "they did not exit after SIGTERM", when an operation
needed SIGKILL. While agent mode shows its progress lines, the line is written among them.
If rushd still runs, it says that only the connection closed. Ctrl+C and an orderly daemon stop
or daemon restart still end a command as cancelled (exit code 130).
A command that was still waiting in rushd's queue when rushd exited has not run, if rushd says
when it starts a command (protocol 0.14) and had not said so. The client then sends it to a new
daemon once, within its --wait-timeout, and before that daemon starts it prints one line on
stderr (or above the agent progress rows): rush-client: rushd (PID <pid>) exited while the command was queued; sending the command to a new daemon. In agent mode the phase, and the status
lines on a pipe, then read request resubmitted to the new daemon; preparing the workspace graph
rather than what the command waited for in the exited daemon's queue, until the new daemon reports
a queue position or starts the command. If the new daemon does not start,
the command fails after that line with the startup error. If the connection to the new daemon
is lost too, the diagnostic begins "Daemon disconnected before delivering a result; the command
was already sent to a new daemon once." Commands that waited together reach the new daemon in
the order in which their clients noticed that rushd exited, not in their order in the queue.
While a command runs, the client checks that rushd still responds. Once rushd has sent nothing
for 10 s, the client pings it. Once it has sent nothing for 30 s, not even the reply, for example
because its process was stopped, the client says so at once and says what that means:
rush-client: rushd (PID <pid>) has not responded for 30s; its process may be stopped or overloaded; on Linux, "rush-client daemon status" says which. This command goes on if rushd responds; interrupt it (Ctrl+C) to stop waiting. When
rushd sends anything again, a second line says so: rush-client: rushd (PID <pid>) responded again after 70s. In agent mode, the progress phase says it; on a pipe both lines are written at once,
and until rushd responds, the status lines say how long it has not responded instead of what runs.
Time in which the client itself was stopped or busy writing output does not count. An interrupt
asks rushd to cancel the command, which a stopped rushd cannot confirm, so the client stops waiting
when its 5 s cancellation wait ends, and writes no more of these lines once it asked. Daemons older
than protocol 0.13 are not checked.
When the process that reads the client's output exits first, for example head in
rush-client build | head -5, the client's next write to that stream fails with EPIPE. The client
does not report that as a lost connection. It asks rushd to cancel the command, waits for the stop
as it does after Ctrl+C, and exits with code 141 (128 + SIGPIPE), which a shell reports for a writer
that SIGPIPE ended. Instead of the cancelling and cancelled lines it prints one line, which names
the stream: rush-client: build cancelled, because the process reading its stdout exited (EPIPE).
In rushx-client it begins with rushx-client:.
Agent output prints the same line on stderr. The client only learns of the exit at its next write,
which in agent output on a pipe can be the next status line, up to 25 s later. A command whose
result arrived before a write failed keeps the result's exit code. Rush run in-process and
rush-client daemon logs still report a failed write as an error.
Piped input uses protocol 0.7's negotiated stdin admission and EOF. The client does
not read input until the command attaches an input destination, and sends bounded
chunks only as the daemon grants write credits. EOF follows all preceding writes;
binary Ctrl+C bytes in a pipe are data, not cancellation signals. Older peers fall
back before requestStart or input consumption, and pre-execution command fallback
preserves the complete pipe for the native entrypoint.
The existing Rush entrypoints resolve project scripts from cwd. Fallback loads the
existing @microsoft/rush version-selecting entrypoint in the client process,
preserving its startup checks, output and reporter integration instead of
inventing a cached module path. Before auto-start or explicit start/restart, the client
selects an available daemon whose installed engine is exactly the requested Rush
version, including a native RUSH_PREVIEW_VERSION override. It can install a published
daemon release declaring that exact engine dependency into a node-specific Rush cache;
it never overrides dependencies or relabels the bundled engine. Foreign installations
are probed in isolation, and startup rechecks actual runtime version, protocol, and
default request-launch APIs before binding. Incompatible or unavailable launchers use
native fallback for ordinary invocations and fail explicitly for management commands.
Connect-only calls never install packages. Older Rush versions may also reject the new daemon config block; use
environment-only opt-in until a supporting Rush release is selected.
Every setting uses environment > config > default. Boolean overrides accept
only 0/1; numeric overrides accept finite unsigned decimal numbers. Unknown
keys and unknown RUSH_DAEMON* variables fail validation.
daemon key |
Environment override | Default | Runtime status |
|---|---|---|---|
enabled |
RUSH_DAEMON |
false | Client routing |
autoStart |
RUSH_DAEMON_AUTO_START |
true | Only after opt-in |
idleTimeoutSeconds |
RUSH_DAEMON_IDLE_TIMEOUT_SECONDS |
900 | Host idle shutdown after request/output/cleanup drain |
queueTimeoutSeconds |
RUSH_DAEMON_QUEUE_TIMEOUT_SECONDS |
30 | Admission wait limit. Time behind another request's graph load (up to 10 times the limit) and the request's own work do not count. The default does not limit waiting behind a running compatible build, or, for a daemon restart, behind requests that were already running (while no rushx script is running); an explicit value does |
watch |
RUSH_DAEMON_WATCH |
false | Persistent host observation of requested warm projects; false keeps root/config guards only. Never schedules builds |
usePersistentIpcRunners |
RUSH_DAEMON_USE_PERSISTENT_IPC_RUNNERS |
false | Enables explicit per-operation daemonIpc Node launchers for unsharded incremental daemon builds |
incrementalBuilds |
RUSH_DAEMON_INCREMENTAL_BUILDS |
true | Runs an operation's <phase>:incremental script instead of its initial script when only files it builds were edited since its last successful run in the daemon and its output folders are unchanged. Additions, deletions, renames, configuration, tool, environment and command-line changes, bundled outputs, cache restores and native Rush commands run the initial script. Incremental results are never written to the build cache |
warmWorkers |
RUSH_DAEMON_WARM_WORKERS |
false | With incrementalBuilds, keeps a watch-mode worker (the <phase>:incremental:ipc script) alive between builds for each operation whose rush-project.json operation settings set allowDaemonWarmWorker, and sends it the next incremental run. Opt in only if that script runs every task and check that the initial script runs, for Heft including lint and API Extractor. When an incremental run is not allowed, the worker is closed and the initial script runs. Workers count toward warmMemoryBudgetMB and warmSetMaxProjects, so raise both to keep them alive |
joinRunningBatch |
RUSH_DAEMON_JOIN_RUNNING_BATCH |
false | Experimental. A build request that arrives while the daemon executes an incremental batch with the same request settings adds its operations to the executing iteration and gets its result once they complete, instead of waiting for the iteration to end. When the iteration can't take its work, the request waits as before |
deferCacheWrites |
RUSH_DAEMON_DEFER_CACHE_WRITES |
false | Lets an operation complete once its output files are cloned, and writes its build cache entry from the clones in the background. Needs a file system that can clone files (such as Btrfs, XFS or APFS); otherwise, and in cobuilds, the entry is written before the operation completes. A failed background write doesn't change the operation's status; the next command reports it in a warning (on stderr in legacy output), which agent output also prints. A restore of the entry's key misses until the entry is written, and daemon stop drops the entries that aren't written yet |
backgroundPrepare |
RUSH_DAEMON_BACKGROUND_PREPARE |
false | Experimental. When a watched workspace input changes so that the next request would reload the workspace graph (lastReloadTier 1), for example after an edit of rush.json or common/config/rush/command-line.json, an idle daemon reloads it and creates the engine for the command line of the last phased command that it served, 2 seconds after the last change. It never runs an operation or restarts the daemon, and it doesn't start while another Rush process holds the repository lock. A request with the same command line waits for it; any other request stops it and runs as before |
warmIdleTimeoutSeconds |
RUSH_DAEMON_WARM_IDLE_TIMEOUT_SECONDS |
300 | Idle runner, project-watcher and retained-result eviction |
warmMemoryBudgetMB |
RUSH_DAEMON_WARM_MEMORY_BUDGET_MB |
512 | Best-effort sampled RSS budget in MiB, not a hard ceiling. Compared against whole-daemon RSS plus measured child RSS, so keep it above the daemon baseline (~130-190 MiB) |
warmSetMaxProjects |
RUSH_DAEMON_WARM_SET_MAX_PROJECTS |
20 | Best-effort limit on projects holding warm resources (active runners, watchers); retained results of resource-free projects do not count. Never trims requested execution |
autoWarmByTelemetry |
RUSH_DAEMON_AUTO_WARM_BY_TELEMETRY |
false | Measured retention ranking with conservative LRU fallback; no speculative scripts |
For genuine persistent Node execution, enable usePersistentIpcRunners and add
operationSettings[].daemonIpc: { entryPoint, args? } in the project's config/rush-project.json.
entryPoint is explicitly project-root-relative, including for inherited/rig settings, and must reside in a
dedicated implementation subdirectory. Node is spawned directly on Linux and Windows; args are literal tokens,
followed by non-ignored native custom parameters. Arbitrary shell strings are not reinterpreted.
The complete implementation directory is fingerprinted (256 entries, 16 levels, 8 MiB maximum); imports outside
it other than Node built-ins are unsupported. Ordinary input/output files must live outside that directory.
Code or descriptor changes replace the old generation, while ordinary input changes reuse the child.
The tool must implement the native IPC protocol and report real RSS; no extra runs manufacture telemetry.
IPC is non-cacheable. Rebuild, NoOp/missing-script behavior, preassigned shards and non-opted-in/native paths
stay unchanged, and existing watch-only IPC declarations do not activate this mode.
Ordinary daemon status reports raw measured workspace.warmSet.projectRanks when available.
Timeouts must be positive and at most 2147483.647 seconds; queue timeout additionally
accepts zero and is rounded down to milliseconds. Memory budget must be positive
and no larger than JavaScript's maximum safe integer. Project count must be a
positive safe integer. The session automatically owns these warm policies for its real
graph and watchers. Executing/prepared work and protected resources are not evicted.
Missing child-memory measurements stay explicitly unknown; unavoidable active/base
memory pressure is reported rather than hidden. No warm-set setting changes build correctness.
Project observation previously ran regardless of watch. Its existing default false
now disables host project observation; set it to true to retain observation between
requests. Every explicit native request still refreshes inputs and effective configuration.
Changing this flag neither discards warm results/runners nor starts scripts; safe idle
maintenance applies watcher changes and reports deferred or failed cleanup.
rush-client daemon start explicitly requests startup, independently of
daemon.enabled, autoStart, or CI execution routing. It conflicts with
--no-daemon. It is idempotent: an existing compatible daemon is reused, not
reconfigured. Startup uses the same detached, locked launcher as automatic
startup and selects an attested launcher rather than guessing a path for another Rush version.
With an explicit matching launcher, a daemon implementation-version mismatch triggers
ownership-checked replacement under the start mutex before executing any command.
It does not replace a peer lacking safe shutdown support. Foreign package installation
is a client preparation step; host self-restart selects only bundled or already cached
compatible installations, never installing while the old workspace is being cleaned up.
Every client of a checkout finds its daemon in one per-user runtime folder: on Linux and
macOS, /tmp/rushd-<uid>/, whatever TMPDIR or XDG_RUNTIME_DIR a shell, job, service or
sandbox sets. It holds the socket (<key>.sock), the ownership record
(<key>.pid.json) and the launcher log. To move it, set RUSHD_RUNTIME_DIR to an absolute
path for every client of that checkout; the folder becomes $RUSHD_RUNTIME_DIR/rushd-<uid>/,
and its file system must support hard links. A relative RUSHD_RUNTIME_DIR is ignored. Clients
that disagree about RUSHD_RUNTIME_DIR use different folders, so each folder gets its own daemon
for the checkout. A client passes the folder to the daemon it starts.
The socket path must fit in a socket address: at most 108 bytes on Linux and 104 on macOS.
RUSHD_RUNTIME_DIR can therefore be at most 57 bytes on Linux and 53 on macOS, minus the
number of digits in your uid (50 bytes on Linux for uid 1234567).
Windows uses the named pipe \\.\pipe\rushd-<key> and is unchanged.
The client refuses a runtime folder that is a symbolic link, is not a directory or belongs to
another user, and a RUSHD_RUNTIME_DIR too long for the socket path: commands run in-process
with that reason, and daemon commands exit 1. Remove the folder or change RUSHD_RUNTIME_DIR.
Auto-start likewise refuses a launcher log that is not a regular file of yours with one link, or
that it cannot open for writing, such as a symlink, a directory or a FIFO: commands run in-process
with that reason (for example Launcher log cannot be opened for writing (ELOOP): <path>), and
daemon start exits 1. Remove the file.
A folder that others can open is made owner-only (0700).
Within one runtime folder, TMPDIR, TMP, TEMP, XDG_RUNTIME_DIR and RUSHD_RUNTIME_DIR
never select a different daemon; each operation receives the requesting client's values.
Clients and daemons before protocol 0.12 used $XDG_RUNTIME_DIR/rushd-<uid>/ or the
temporary folder instead. A daemon started there stays there, where current clients do not
look, until it idles out or is stopped with that older client (rush-client daemon stop).
An older client that starts a current engine while XDG_RUNTIME_DIR or TMPDIR is set does
not find it and runs in-process, so upgrade rush-cli-client with the engine.
When a daemon before protocol 0.12 serves a current client, for example one listening in
/tmp/rushd-<uid>/, the client leaves XDG_RUNTIME_DIR, TMPDIR, TMP and TEMP out of
its requests on Linux and macOS, because that daemon would restart into the folder they name.
Its operations see the daemon's own values; stop it (rush-client daemon stop) to use yours.
rush-client daemon status only connects and checks hello/pong. It never starts
a process, reclaims files, or treats a PID file as evidence of readiness. Both
commands print one JSON object with state: "ready", socketPath, and the actual
pong fields (uptimeMs, available versions, optional pid and
residentMemoryBytes, and an optional workspace snapshot). Exit code 0 means protocol
readiness, not build support. An unreachable/incompatible endpoint, invalid
arguments, or startup failure returns exit code 1 with a diagnostic. When no daemon runs
at all, status also exits 1, and its diagnostic says No daemon is running for <repository> (Rush <version>) and what the next command does: with enabled and autoStart, the next
rush-client command that uses the daemon starts one; otherwise rush-client commands run Rush
in-process. This is the normal state after daemon stop, the idle timeout or SIGTERM. Status
reports it only when the endpoint refuses connections and neither an ownership record, a startup
reservation nor (outside Windows) a socket file remains; with any of these, the diagnostic
still says that it could not connect. When the endpoint
refuses connections and its ownership record (<key>.pid.json) names a PID that no longer
exists, the diagnostic adds that rushd exited without shutting down (an orderly shutdown
removes the record) and that daemon logs may show why. When that PID still exists but the endpoint
refuses connections or does not complete hello/ping, the diagnostic adds what that process is doing,
on Linux for example rushd (PID <pid>) still owns <key>.pid.json: it is stopped (state T), for example by SIGSTOP, and it started 5 min ago., and on the next line what to do about it.
A client that lost its connection to that daemon, or that ran Rush in-process, removes the record
when it reclaims the daemon, and appends a line that names the daemon to the launcher log. Until a
daemon becomes ready again or daemon stop --force resets the workspace, the No daemon is running
diagnostic then adds The last daemon, rushd (PID <pid>), exited without shutting down; "rush-client daemon logs" may show why.
A daemon whose installation was removed or replaced still answers, but it restarts on
the next command: status then prints state: "installationChanged" with the pong's
installationChange (change and folder), a hint on stderr, and exits with code 1.
A startup reservation (<key>.pid.json.starting) refuses another daemon launch until
the daemon it reserved becomes ready. Status reports one that remains as
startupReservation with its path, the startup helper's helperPid when recorded, and
helperState: running (the helper still waits for readiness), exited (the helper will not
release it), or unknown (written by an older client). Status never removes it. Next to a
ready daemon, the next command that uses, stops or restarts that daemon removes it; when status
cannot connect, its diagnostic explains the reservation. After an exited helper, status also
reports relaunchAfter, 15 seconds after that helper was launched. Until then every automatic
start is refused at once (the command runs in-process), so that a daemon that fails the same way
each time, for example because of a configuration error, is not launched by every command. The
first command after it that finds nothing listening at the endpoint takes the reservation over
and starts the daemon again, and daemon logs shows a line saying so. daemon logs may also show
why the daemon did not become ready.
The optional workspace snapshot reports the provider generation/token, graph existence, and available warm accounting without initializing a graph. Missing fields are unknown, not proof of zero memory or successful reload. Status can inspect a protocol-compatible daemon with a different implementation version; start requires the bundled version to match.
workspace field |
Meaning |
|---|---|
generation, generationToken |
Current provider generation and installed session identity |
lastReloadTier |
Lifecycle-owned 0 initial/reuse, 1 successful in-process reload, or 2 requested restart; older peers may omit it |
graphInitialized |
A graph exists; this does not attest build success |
warmSet.configuration |
Effective watch and four warm-resource settings; older peers may omit watch |
warmSet.maintenanceState, warmSet.maintenanceFailure |
Running, quiescing, stopped or failed maintenance; stopping it does not itself free resources |
warmSet.retainedProjectNames, warmSet.protectedProjectNames, warmSet.watchedProjectNames |
Actual retained/protected projects and resident project observation |
warmSet.daemonResidentMemoryBytes, warmSet.measuredRunnerMemoryBytes, warmSet.unmeasuredRunnerCount |
Daemon RSS, last-completion child RSS samples, and explicitly unmeasured resident runners; descendants are not included |
warmSet.overMemoryBudget, warmSet.overProjectLimit, warmSet.cleanupFailures, warmSet.deferredReason |
Outstanding footprint pressure, cleanup failures and maintenance deferral |
An absent warmSet means no controller is attached, not that the workspace consumes
no memory. Status reads lastReloadTier from the lifecycle (zero for a host without one);
it does not infer a tier from PID/generation changes or initiate a reload. Tier 2
attests a restart request, not completion of successor startup or success of a command.
rush-client daemon stop requires protocol >= 0.6 and waits for shutdownAck
followed by EOF. It reports state: "shutdownAccepted" with exit code 0; this
does not assert successful workspace disposal. Stop is idempotent: when nothing
listens at the endpoint and no daemon is starting, it reports state: "notRunning" with exit code 0.
When nothing listens but the ownership record names a process that still runs, for example a daemon
that removed its socket while it shuts down, stop first waits up to 15 seconds for that process to exit,
and says so on stderr once it has waited a second. If it still runs then, stop exits with code 1 and
says what that process is doing and what to do about it; a daemon that accepts the connection but does
not complete hello/ping gets the same diagnostic. On Linux, stop does not wait for a Rush daemon that
has this workspace's ownership record open and stays stopped (state T or t) while stop samples it for
1.5 s, because it cannot exit before something resumes it: stop exits with code 1 and that diagnostic
once the 1.5 s have passed.
While a daemon is still starting (its startup helper still runs, or another client holds the start
mutex), stop says so on stderr and waits up to 15 seconds for that daemon to become ready, then stops
it as below; reporting notRunning would leave it running afterwards. If it is still not ready by then,
stop exits with code 1 and leaves it running; run stop again once daemon status reports it ready. An
unsupported protocol, missing acknowledgement, handshake failure, or timeout
returns exit code 1. It does not auto-start anything. Before shutdown, it removes a startup
reservation that remains next to that daemon, as restart does, so that the reservation cannot
refuse the next start once the daemon is gone. It does so only for the live owner in the
ownership record, under the start mutex (waiting up to 15 seconds for it); a reservation that
it cannot resolve stays in place and is reported as startupReservation.
rush-client daemon stop --force stops a running daemon the same way, then waits
(up to 15 seconds) for it to release its listener and ownership record and removes
any remaining artifacts, such as an abandoned startup reservation, reporting them in
removedPaths. When none is listening, it removes this workspace's leftover ownership record
(<key>.pid.json), socket, and startup reservation (.starting), then reports
state: "reset" and the removedPaths (or state: "notRunning" if nothing was
left behind). It holds the start mutex, proves that no listener is bound, and
refuses (exit 1) while the recorded owner PID still exists and cannot be shown to
be a reused PID, saying what that process is doing, that no process was killed, and what to do about it.
It fails the same way, without killing that process, when that process accepts the connection but does
not complete hello/ping, for example because a signal stopped it. When the recorded owner PID no longer exists, the daemon exited without shutting
down and may have left operations running that only its records name, so the reset first stops them
as the next daemon start would: SIGTERM, then SIGKILL 2 seconds later, to the daemon's own process
group and to each operation process group that it recorded whose leader still has the recorded start
time (or has exited, while every live member of the group is in the group's own session and one of
them still has the RUSHD_OPERATION_GROUPS variable that the daemon gives the processes it starts). When a
process that started after the record was written has the recorded PID now, the daemon exited the same
way, so the reset stops the operation process groups that it recorded the same way, but never the
process group whose ID is that PID, which the later process may lead. It prints the
line shown above for a lost connection and reports what it stopped in orphansReaped (daemonPid,
processGroupIds, outcome). A recorded group that it cannot prove, such as a PID that a later
process now has, or a group whose leader has exited under a daemon from a release that did not set
RUSHD_OPERATION_GROUPS, is not signalled; its record is removed with the others. When such a group still
has a live process, the launcher log that daemon logs prints gets a line that names the group and the
daemon and says which check the group failed; nothing about it is printed. The reclaims after a lost
connection, before Rush runs in-process and before a daemon start do the same. If the operations cannot be
stopped, it exits with code 1 and removes nothing. While another process reclaims the same files it
also exits with code 1, except that after a shutdown it re-checks for up to 15 seconds. Otherwise it
never signals a process. Automatic startup already reclaims
the common leftovers on its own (see below); this is the documented escape hatch
that every fail-closed startup message points to. A reset also appends a line to the launcher log
that clears the report of a daemon that a client reclaimed, so status no longer names it.
rush-client daemon restart first verifies that the selected Rush version has a
launcher and captures the original lock's PID/start timestamp, checking that it
matches pong's positive PID and the selected endpoint, then performs acknowledged
shutdown. Before shutdown, it removes a startup reservation that remains next to that
daemon (under the start mutex), so that the reservation cannot refuse the successor; if
another client holds the mutex for 15 seconds, restart fails without stopping the daemon.
It waits for original ownership release or a demonstrably dead owner
before calling the existing locked starter. A live owner fails closed at
the startup deadline; no PID is killed and no live ownership record is deleted.
A newly
started/reused successor must pass hello/ping before reporting state: "ready".
When nothing listens at the endpoint, restart starts a daemon exactly like daemon start.
Automatic and explicit startup reclaim stale artifacts only when that is provably
safe: while holding the start mutex with no .starting reservation (or after taking over
one whose helper exited, as described above), a socket
without an ownership record, or an unreadable/corrupt record, is removed only after
a connection attempt is refused (so no listener exists). On Linux, a record whose
PID now belongs to a process that started after the record's startedAt (PID reuse)
is treated as dead; other platforms fail closed and point to daemon stop --force. Before such a
record is removed, the operation process groups that the daemon recorded are stopped as
daemon stop --force stops them, and the line shown above for a lost connection is printed; if they
cannot be stopped, the command fails and the record stays.
Restart is explicit even when automatic startup or CI execution routing is
disabled, but conflicts with --no-daemon. The two-phase host retains ownership
until workspace disposal succeeds, so embedded hosts can restart a workspace
without exiting their process. Failed cleanup retains the live lock and causes a
bounded restart failure, even if the socket has already disappeared. A changed
owner is reconnected and validated, not overwritten.
rush-client daemon logs prints a snapshot of the selected workspace's launcher
log, whether the daemon is running or stopped. It never connects or auto-starts.
The stable path comes from getDaemonLogFilePath(paths): <lockfilePath>.log.
Detached child stdout and stderr are appended to this file across restarts; the
parent closes its descriptor after spawning. On POSIX the launcher enforces mode
0600 and rejects linked destinations; Windows uses the existing per-user
transport directory permissions.
Each daemon writes <time> rushd ready at <socket> (Rush <version>, PID <pid>) when it
is ready, and <time> rushd (PID <pid>) shutting down: <reason> when it begins to shut
down. The reason is received SIGTERM or received SIGINT; requested by a client, for
daemon stop, daemon stop --force and daemon restart; idle for <seconds> s;
its socket file was deleted or replaced, after the line that reports the lost socket;
or a restart, with its cause. A daemon that is killed (SIGKILL, or SIGHUP, which it does
not handle) writes no such line, and a crash still writes its stack.
Default reading is bounded to the size observed when the log is opened, with chunked, backpressured output. Empty snapshots succeed without output; missing/unreadable logs or invalid destinations fail with a diagnostic and exit code 1. A launcher log may not exist for a daemon started outside this client.
rush-client daemon logs --follow also streams bytes appended after EOF, using
at most 64 KiB per read and waiting for output backpressure before reading more.
It never connects or starts a daemon, including when following an empty log.
SIGINT/SIGTERM cancel following with exit code 130 after closing the log descriptor;
pending display bytes may be discarded rather than keeping the client alive
indefinitely behind a stalled output pipe. Read/write failures remain explicit errors. Truncation,
replacement, or removal fails explicitly instead of silently following the wrong
file; reopen the command after rotating a log. No automatic rotation policy is added.
On Windows, following redirected stdout uses an invocation-owned, output-only Node
process. Windows pipe writes can block the CLI event loop, and moving a blocked
write to its filesystem thread pool alone can still prevent shutdown. The CLI
instead sends one acknowledged chunk of at most 64 KiB to the isolated writer.
Cancellation stops and joins that writer before the CLI exits; a lost parent
also terminates the writer, so blocked display output is not orphaned. The worker
does not run user NODE_OPTIONS preload hooks or any Rush command. This is not
daemon startup. Default snapshots, TTY output, and the non-Windows output path
retain their existing stdio behavior.
This is the text launcher stdout/stderr log, including startup errors—not WS5 structured observability or a subscription to request-scoped events.
Set RUSH_DAEMON_EXPERIMENTAL=1 and use an explicitly started daemon:
export RUSH_DAEMON_EXPERIMENTAL=1
rush-client daemon start
rush-client daemon graph show
rush-client build --to my-project
rush-client daemon graph scope-out --project my-project
rush-client daemon graph scope-in --operation 'my-project (compile)'
rush-client daemon graph invalidate --project my-project
rush-client daemon graph pause
rush-client daemon graph status
rush-client daemon graph resume
rush-client daemon graph watchAll graph commands connect only; they never auto-start, initialize the graph, or
fall back to native Rush. Cold show and status report initialized: false
without an operations field. Other verbs fail explicitly until a supported,
explicit build request has initialized the graph. The gate is checked both by the
CLI and against the server request's environment. Older or unsupported servers,
invalid arguments, and unknown selectors fail, never invoke a shell.
show and status both emit a complete point-in-time metadata snapshot, including
operation IDs, exact project/phase names, native enabled states, observed statuses,
dependency IDs, manual-mode/scheduled flags, and a path-free invalidation summary.
An operation without an observed execution status reports null. An idle operation
whose actual completed result was evicted reports READY for request-time revalidation,
not historical success; inspection does not schedule work. Snapshots contain no
environment, runner, log, or terminal objects.
Protocol 0.9 snapshots include an opaque workspaceGeneration token. Every mutation
echoes a token, checked under exclusive admission before touching the graph. The
token changes on soft reload and process replacement, preventing stale operation
references from affecting a new generation. Use --generation TOKEN with a token
from an earlier snapshot to preserve that reference; the client never refreshes an
explicit token. Without this option, the client privately reads current status
before submitting the mutation. A reload between those requests fails closed.
Mutations reject older peers before submission; read-only inspection remains compatible.
Graph requests use the same resolved queue options as ordinary execution:
--no-wait/--wait-timeout override environment, config, and the 30-second default.
An implicit generation lookup and its mutation share one positive admission
deadline; time spent connecting and reading the token is not reset before the
mutation. If the budget expires, no later query is sent. Zero requests immediate
server admission, as does --no-wait, without imposing a zero-length connection
timeout. Cancellation covers generation preflight too. Once a mutation is admitted,
its execution is not timed out by the admission budget.
scope-in, scope-out, and invalidate require one or more repeated
--project NAME or --operation ID pairs. Names and IDs match exactly; there are
no globs or implicit all-project selections. Every selector is validated before
any mutation. Scope-in enables transitive dependencies. Scope-out includes
transitive consumers and uses native safe-disable, which also prunes dependencies
no longer needed by enabled operations. Invalidation marks selected native results
stale without scheduling work. Mutations use exclusive workspace admission and
never change an active iteration. Scope changes/invalidation reject an already
prepared iteration instead of modifying stale execution records.
Pause/resume change native pauseNextIteration: manual mode gates automatically
scheduled iterations, not explicit build requests. Resume does not create new
work. If an engine owner has already prepared an automatic iteration, resume
acquires the native execution lease, discards the old unstarted plan, reconciles
inputs, and prepares its replacement before releasing it. Admission and the native
lease remain held until native idle, including
if the resume client disconnects. The reference client's watch command is not a
build scheduler; the default lazy engine still rebuilds only on explicit requests.
Native build requests apply their own selections, so a graph scope is not a
persistent override of later build arguments.
Graph stdout is NDJSON only. Snapshots use the existing extension event envelope
with payload.name: "rushd.graph-snapshot" and
payload.data: { requestId, snapshot }. The existing requestResult is emitted as
the terminal record; queue positions and request rejections also retain their
control-message shapes. Local failures use { kind: "graphError", message }.
No human renderer, ANSI styling, or graph-specific transport is involved.
watch emits an initial snapshot followed by relevant graph state, invalidation,
and idle updates. It holds no scheduler lease, so other clients can build.
Slow consumers receive coalesced latest snapshots rather than every intermediate
transition or an unbounded event history. SIGINT/SIGTERM cancel the subscription
and wait for the authoritative aborted result (exit 130). Disconnect removes
subscriber resources without cancelling another client's build. Graph hooks are
installed once per graph, not once per connection.
Origin issues #5894 and #5896 still require reconciliation with the canonical
reporter contract and built-in reporters. The current wire consumers use
IDaemonEventEnvelope/DaemonEventType; ClientOperationRenderer feeds
DaemonRendererHost, whose default is LegacyCollatedRenderer. These inherited
contracts are not replaced by the WS4 terminal, graph-admission, or log-follow
fixes. Explicit reporter controls continue to select the native reporter path.
Reporter reconciliation remains a separate epic decision, not an assertion that
the WS5 default flip or the full daemon transparency matrix has passed.