Description
There is a race condition in startGoRoutines (functions/daemon.go) that causes
servers.json and nodes.json to be permanently wiped to {}, leaving the daemon
alive but completely blind - all subsequent netclient list, netclient ping, and
netclient pull calls fail with "server config not found".
This is distinct from the directory-deletion race fixed in v1.5.0 that we raised previously (#1186)
How It Happens
The race requires two things in tight sequence:
WriteJSONAtomic is mid-write on servers.json - specifically between step 2
(rename servers.json → servers.json.bak) and step 3
(rename servers.json.tmp → servers.json)
daemon.Restart() (SIGHUP) fires - triggered by any MQTT peer/host update
When SIGHUP fires in that window:
- New daemon starts, calls
ReadServerConf() - but servers.json doesn't exist yet
(renamed to .bak, .tmp not yet renamed to final)
ReadServerConf returns error, Servers stays as the empty map from init()
startGoRoutines proceeds normally with empty Servers
- Next MQTT message triggers
WriteServerConfig() with the empty in-memory map
→ writes {} to disk permanently
After this point, netclient pull cannot self-heal because Pull() in
functions/pull.go reads config.GetServer(serverName) from in-memory first -
which is already empty - and returns "server config not found" before making
any HTTP call. Recovery requires a full process restart.
Evidence
Observed on v1.5.0. Container ran healthy for ~1h40m after enrollment, then:
# /etc/netclient/ state after the race:
-rwx------ servers.json # content: {} (wiped)
-rwx------ nodes.json # content: {} (wiped)
-rwx------ servers.json.bak # content: full valid config (from enrollment)
-rwx------ nodes.json.bak # content: full valid config (from enrollment)
# All netclient commands fail:
$ netclient list
no such network
$ netclient ping
Failed to ping peers: server config not found
$ netclient pull
fail to pull config from server: server config not found
Affected Files
Both servers.json and nodes.json are vulnerable - they follow the identical
pattern in config/server.go and config/node.go:
ReadServerConf / ReadNodeConfig fail silently on missing file
- In-memory map stays as empty
make(map...) from init()
- Next write call flushes empty map to disk
Both need the fix.
Proposed Fix
Add a backup restore helper to config/server.go:
func RestoreServerConfFromBackup() error {
bakFile := filepath.Join(GetNetclientPath(), "servers.json.bak")
liveFile := filepath.Join(GetNetclientPath(), "servers.json")
data, err := os.ReadFile(bakFile)
if err != nil || string(data) == "{}" || len(data) < 5 {
return fmt.Errorf("no valid backup")
}
return os.WriteFile(liveFile, data, 0700)
}
Same pattern for config/node.go (RestoreNodeConfFromBackup).
Then in startGoRoutines (functions/daemon.go), after each ReadServerConf
call (there are two call sites, lines ~156 and ~254):
if err := config.ReadServerConf(); err != nil {
slog.Warn("error reading server map from disk", "error", err)
}
// TOCTOU recovery: restore from backup if Servers is empty after read
if len(config.GetServers()) == 0 {
if err := config.RestoreServerConfFromBackup(); err == nil {
slog.Info("restored server config from backup after TOCTOU race")
_ = config.ReadServerConf()
}
}
Apply the same pattern for ReadNodeConfig / RestoreNodeConfFromBackup /
len(config.GetNodes()) == 0.
Additional Context
We discovered this while running agents on edge hardware where MQTT peer/host
updates are frequent (multiple nodes joining/leaving the same cluster network).
The trigger in our case was an MQTT-driven daemon restart ~1h40m after startup,
which happened to land in the WriteJSONAtomic race window.
Description
There is a race condition in
startGoRoutines(functions/daemon.go) that causesservers.jsonandnodes.jsonto be permanently wiped to{}, leaving the daemonalive but completely blind - all subsequent
netclient list,netclient ping, andnetclient pullcalls fail with"server config not found".This is distinct from the directory-deletion race fixed in v1.5.0 that we raised previously (#1186)
How It Happens
The race requires two things in tight sequence:
WriteJSONAtomicis mid-write onservers.json- specifically between step 2(rename
servers.json→servers.json.bak) and step 3(rename
servers.json.tmp→servers.json)daemon.Restart()(SIGHUP) fires - triggered by any MQTT peer/host updateWhen SIGHUP fires in that window:
ReadServerConf()- butservers.jsondoesn't exist yet(renamed to
.bak,.tmpnot yet renamed to final)ReadServerConfreturns error,Serversstays as the empty map frominit()startGoRoutinesproceeds normally with emptyServersWriteServerConfig()with the empty in-memory map→ writes
{}to disk permanentlyAfter this point,
netclient pullcannot self-heal becausePull()infunctions/pull.goreadsconfig.GetServer(serverName)from in-memory first -which is already empty - and returns
"server config not found"before makingany HTTP call. Recovery requires a full process restart.
Evidence
Observed on v1.5.0. Container ran healthy for ~1h40m after enrollment, then:
Affected Files
Both
servers.jsonandnodes.jsonare vulnerable - they follow the identicalpattern in
config/server.goandconfig/node.go:ReadServerConf/ReadNodeConfigfail silently on missing filemake(map...)frominit()Both need the fix.
Proposed Fix
Add a backup restore helper to
config/server.go:Same pattern for
config/node.go(RestoreNodeConfFromBackup).Then in
startGoRoutines(functions/daemon.go), after eachReadServerConfcall (there are two call sites, lines ~156 and ~254):
Apply the same pattern for
ReadNodeConfig/RestoreNodeConfFromBackup/len(config.GetNodes()) == 0.Additional Context
We discovered this while running agents on edge hardware where MQTT peer/host
updates are frequent (multiple nodes joining/leaving the same cluster network).
The trigger in our case was an MQTT-driven daemon restart ~1h40m after startup,
which happened to land in the WriteJSONAtomic race window.