The Linux kernel's implementation of Filesystem in Userspace
requires root permissions, despite its use in unprivileged programs.
This has normally been solved via libfuse's setuid helper program
fusermount/fusermount3.
This does means that certain kinds of security policies cannot be applied,
specifically no_new_privileges process flag.
$ mkdir _lower _mnt
$ # Without no_new_privileges
$ fuse-overlayfs -o lowerdir=_lower _mnt
$ fusermount3 -u _mnt
$ # With no_new_privileges
$ setpriv --no-new-privs -- fuse-overlayfs -o lowerdir=_lower _mnt
/usr/bin/fusermount3: mount failed: Operation not permitted
fuse-overlayfs: cannot mount: Operation not permittedThe no_new_privileges flag is important for proper application sandboxing,
as Linux features such as landlock and seccomp-bpf can only be used
after a call to prctl(PR_SET_NO_NEW_PRIVS, 1).
By using Unix domain sockets, the privileged mounting operation is performed
by a process outside the unprivileged process tree, bypassing the
no_new_privileges restrictions.
Caution
defused is a pre-alpha project: use at your own risk!
Defused requires Linux 6.12 or later as it uses name_to_handle_at()'s
AT_HANDLE_MNT_ID_UNIQUE argument to obtain a unique mount id.
It also depends on libseccomp.
This project provides the following:
- A system service that listens on
/run/defused/defused.sock. - A replacement
fusermount3andfusermountbinary to communicate with the service. mount.fuseandumount.fusehelper scripts formountandumount, sofuse.*filesystems in/etc/fstabwork without libfuse'smount.fuse3, andumountunmounts throughfusermount3.
The system service is written to use systemd socket activation with
Accept=yes.
For testing or on systems without systemd, defused --daemon can be used
to create the socket and fork off child processes to handle accepted
connections.
Instead of a config file (like libfuse's /etc/fuse.conf), all of defused's
options are configured on the daemon's command line:
| Option | Meaning | Default |
|---|---|---|
--max-mounts=N |
Refuse a mount once N FUSE filesystems are mounted (libfuse's mount_max) |
100 |
--allow-groups=GROUP[,GROUP...] |
Only members of these groups (names or gids) may mount and unmount | any user |
--allow-other |
Let callers set the allow_other mount option (libfuse's user_allow_other) |
refused |
The ownership checks below always apply as well.
The installed defused@.service takes these from environment variables, so
changing them is a drop-in:
# systemctl edit defused@.service
[Service]
Environment=DEFUSED_ALLOW_GROUPS=fuse
DEFUSED_MAX_MOUNTS works the same way. DEFUSED_EXTRA_ARGS is split on
whitespace and appended to the command line, which is how a flag-only option
like --allow-other is passed:
# systemctl edit defused@.service
[Service]
Environment=DEFUSED_EXTRA_ARGS=--allow-other
A caller that is root or holds CAP_SYS_ADMIN does not need the service at
all, so fusermount3 will bypass the socket in that case and perform the
mount directly.
Defused uses a different mountpoint ownership model than libfuse's setuid
fusermount3.
For non-root mounts, the mountpoint must be a directory or regular file owned
by the caller.
It must be writable by that caller, and directories must also be searchable.
This means defused rejects mounts on writable shared directories owned by
another user, even when libfuse's setuid helper would allow them because the
directory is not sticky.
The stricter rule keeps the privileged service's authorization decision tied
to the mountpoint file descriptor it receives, instead of trying to reproduce
libfuse's path-based access(W_OK) check across the client/service protocol.
This does lead to some additional mounting possibilities, all due to other
filesystem restrictions.
If a given file path is owned by the user, but the process is unable to write
to the path due to POSIX ACLs, LSMs like SELinux, AppArmor, or Landlock,
libfuse's setuid implementation will deny the mount while this implementation
will still perform it.
I do not believe this is an issue, however, as sandboxed applications should
deny access to /dev/fuse or /run/defused/defused.sock.
Otherwise you could use systemd-run --user to "escalate to user".
See protocol.md for more information on how defused works.
{
imports = [ inputs.defused.nixosModules.defused ];
services.defused.enable = true;
}This replaces /run/wrappers/bin/fusermount3 and /run/wrappers/bin/fusermount
with defused's, so every FUSE program uses it, libfuse2 and libfuse3 alike.
services.defused.replaceFusermount3 and replaceFusermount turn either
takeover off.
services.defused.maxMounts, allowGroups and allowOther configure the
policy.
See services.defused.* for the options.
I am using cachix as a binary cache:
# Add to nix.conf
extra-substituters = https://defused.cachix.org
extra-trusted-public-keys = defused.cachix.org-1:/YD+2Bmle49JSliBhGRqTKpLYhvruoFyMPPU071YCAY=
See contributing.md.
The mountpoint filesystem allowlist in defused.c is copied from libfuse and is GPL-2.0-only (marked via an SPDX snippet). All of my code is licensed under GPL-2.0-or-later, but the resulting binary will be GPL-2.0-only.