This is a display of mostly-automatically-classified git commits from 2026-07-01 to 2026-09-30.
Table of contents and commits per category:
| (49) | Highlighted commits (these are copies, not in stats) | |
| 47 | 1.8% | Userland programs |
| 192 | 7.3% | Documentation |
| 937 | 35.5% | Hardware support |
| 192 | 7.3% | Networking |
| 245 | 9.3% | System administration |
| 124 | 4.7% | Libraries |
| 50 | 1.9% | Filesystems |
| 466 | 17.7% | Kernel |
| 72 | 2.7% | Build system |
| 49 | 1.9% | Internal organizational stuff |
| 96 | 3.6% | Testing |
| 49 | 1.9% | Style, typos, and comments |
| 93 | 3.5% | Contrib code |
| 25 | 0.9% | Reverted commits |
| 0 | 0.0% | Unclassified commits |
| 2638 | 100% | total |
| Technical notes about this page |
For extra visibility, these are copies of commits found in
other sections. Most (if not all) come from the commit message
containing "Relnotes:", or commits modifying
UPDATING.
TLS receive offload is really only beneficial for in-kernel use cases (such as NFS over TLS) or when using a hardware offload. In addition, several recent SAs have involved the TLS receive path, but the only current mitigation for those is to disable TLS offload entirely. Reviewed by: ziaee, gallatin, markj Relnotes: yes Sponsored by: Netflix Sponsored by: Chelsio Communications Co-authored-by: John Baldwin <jhb@FreeBSD.org> Differential Revision: https://reviews.freebsd.org/D57974
Start the loop by finding the end of the option name, the name-value separator (if any), and the end of the option. Use those pointers to simplify matching the option name and parsing the option value, and validate option names and values more strictly. This means that: * We no longer accept trailing garbage in an option name or value. For instance, we would previously interpret “edns0123” as “edns0” and “timeout:3xyz” as “timeout:3”. This was actually quite lucky because we also failed to recognize the newline at the end of the option line as a whitespace character. * For options that take a numerical argument, we would previously accept negative values and treat non-numerical arguments as 0, while large numerical arguments would be capped to the option's maximum permitted value. Now, any failure to parse the argument, including overflow, results in the option being left unchanged. MFC after: 1 week Relnotes: yes Reviewed by: markj Differential Revision: https://reviews.freebsd.org/D57923
When we switched from the BIND4 resolver to the BIND9 resolver, the sortlist parser was inadvertently disabled due to a missing #define, and nobody seemed to notice. The sorting code remained enabled in the resolver, but there was no way to set a sort order. Reimplement the sortlist parser, but correctly, and update the manual accordingly. The new parser accepts IPv4 and IPv6 addresses with or without a mask or prefix length, just like the old one, except IPv6 support was a bit wonky in the original code. Fixes: https://cgit.freebsd.org/src/commit/?id=5342d17f09a8 ("Update the resolver in libc to BIND9's one.") Relnotes: yes Reviewed by: markj Differential Revision: https://reviews.freebsd.org/D57925
If a nhop gets an interface event, revalidate the nhops and immediately try to recompile existing nexthop groups by replacing unreachable nexthops with reachable ones. If none are available, recompile them back to their normal position in nexthop group slots. Reviewed by: glebius Discussed with: markj Relnotes: yes Differential Revision: https://reviews.freebsd.org/D57389
The tcp_bblog facility provides structured logging of TCP stack activity for debugging and performance analysis. It is implemented in the kernel and allows per-connection tracing of TCP events with low overhead. Reviewed by: tuexen, ziaee MFC after: 1 week Relnotes: yes Differential Revision: https://reviews.freebsd.org/D56252
This internet draft (which is close to being an RFC) specifies a new NFSv4.2 attribute which tells the NFSv4.2 client to not cache file data. (Similar to O_DIRECT, but triggered by this attribute set on the file on the NFSv4.2 server and not by the application's open(2).) https://datatracker.ietf.org/doc/draft-ietf-nfsv4-uncacheable-files/ This patch adds a new chflags(1) flag called UF_DONTCACHE to implement this. Patches for NFS and ZFS will be done separately. Reviewed by: kib MFC after: 2 weeks Differential Revision: https://reviews.freebsd.org/D58181
Add the option "oemstring" to allow setting the DMI type 11 ("OEM
Strings") SMBIOS structure. These are free-form strings, available for
any purpose, but can be especially useful to pass configuration,
secrets, and credential information into a Linux guest and consumed by
systemd.
MFC after: 1 month
Relnotes: yes
Reviewed by: markj
Differential Revision: https://reviews.freebsd.org/D57516
Point out which features are non-POSIX and thus can not be safely assumed to be portable and exist in other implementations. Relnotes: YES! Reviewed by: ziaee, jilles Differential Revision: https://reviews.freebsd.org/D55333
Use pwait's new -r option to wait until the target processes have not only terminated, but also been reaped. PR: https://bugs.freebsd.org/bugzilla/show_bug.cgi?id=293183 MFC after: 1 week Sponsored by: Klara, Inc. Sponsored by: NetApp, Inc. Reviewed by: markj Differential Revision: https://reviews.freebsd.org/D58391
Add the ability to select source ip address of outgoing packets even when the source ip address is configured on another interface. Also add this new rtnetlink attribute to manual. PR: https://bugs.freebsd.org/bugzilla/show_bug.cgi?id=285422 Reviewed by: glebius, ziaee (manpages) Tested by: ivy, Marek Zarychta <zarychtam@plan-b.pwste.edu.pl> Relnotes: yes Differential Revision: https://reviews.freebsd.org/D58294
Register the 82576 and I350 VF PCI IDs under a separate igbv driver while continuing to share the igb datapath implementation. Follow the ixv driver split and give the VF context IFLIB_IS_VF so iflib does not apply the PF SR-IOV detach guard to a child VF. Program VTIVAR_MISC in the VF low byte so mailbox and reset notifications reach the VF admin vector. The split will become increasingly obvious as bug fixes land, trying to bias everything with if (sc->vf_ifp) everywhere is error prone in two directions. This breaks existing naming/configurations and cannot be MFCed as-is. I have no plans of adapting it to prior branches at the moment but it may be possible. Relnotes: yes Sponsored by: BBOX.io
Add the PCI IOV schema and PF control plane for up to seven VFs with one hardware queue per pool. Implement VF mailbox handling, MAC and VLAN assignment, multicast filtering, promiscuity policy, anti-spoofing, malicious-driver recovery, reset replay, and queue lifecycle management. The basic SR-IOV and VMDq PF implementation follows DPDK Intel e1000 code, including PF pool selection, one queue per pool, mailbox dispatch, and VF enablement. Intel FreeBSD igb-2.5.31 supplies the older driver baseline. Linux igb and the Intel SDMs clear up lifecycle, isolation, reset, and family-specific details absent from DPDK. Enabling IOV requires the PF to attach with one TX and RX queue. Systems whose defaults select RSS queues must set the documented iflib queue override tunables before attach. Only 82576 and I350 support SR-IOV in silicon. The series has been extensively tested on I350, including thowing boundaries at the PCI BAR that shipping drivers will never. Still, think carefully before reaching for this in critical environments. Relnotes: yes Sponsored by: BBOX.io
Document supported controllers, PF and VF naming, PCI_IOV and IOMMU requirements, queue and lifecycle constraints, iovctl schema, filtering and anti-spoof policy, mailbox and MDD recovery, shared hardware limits, rate control, and statistics cadence. Relnotes: yes Sponsored by: BBOX.io
Borrow the e1000 VLAN filter table Ambiguous presence of the feature by Intel was settled by DPDK and emperical testing. MFC after: 2 weeks Relnotes: yes
The SR-IOV schema advertises MAC anti-spoofing and enables it by default, but the VF configuration was never consumed and the hardware policy remained disabled. Record the configured policy and apply MAC and VLAN anti-spoofing throughout VF initialization and reset. On X550-family devices, also protect the LLDP and flow-control Ethertypes and enable per-VF spoof-event accounting. Remove the driver-owned state during SR-IOV teardown. Adapt the anti-spoof configuration lifecycle used by igb(4) in a2ed165f0049 to the ixgbe hardware controls. MFC after: 1 week Relnotes: yes
The VF VLAN capability is checked but never granted, and no SR-IOV configuration property exposes the existing default-VLAN support. PF VLAN updates also replace VFTA registers from a PF-only shadow, erasing live VF filters. Expose access VLAN and trunk policy through the IOV schema. Track each VF VLAN as desired state, restore the administrative VLAN after reset, and use the native VLVF helper for incremental PF and VF ownership changes. Keep VLAN filtering enabled while SR-IOV is active. When PF hardware filtering is disabled, admit every VLAN to the PF without bypassing per-pool VF isolation. Reconstruct VLVF and the shared VFTA from PF and VF desired state after reset or a filtering-mode transition, and restore PF-only state on teardown. When the last VF leaves a VLAN still owned by the PF, free its VLVF slot while retaining the shared VFTA bit. This prevents a trunk VF from exhausting the 64-entry VLVF table by cycling VLAN memberships. Adapt the VLAN ownership model introduced for igb(4) in a2ed165f0049 to ixgbe's native VLVF machinery. Match Linux receive semantics by exposing a stripped VLAN tag only when that VID was registered by the VF. A PF-assigned port VLAN is an administrative tag and must be delivered to the VF as untagged traffic; otherwise the stack dispatches it to a nonexistent VLAN interface and access-VLAN receive traffic is blackholed. MFC after: 1 week Relnotes: yes
The allow-promisc IOV property is advertised but ignored, and the PF rejects the xcast request used by modern VFs. Negotiate mailbox APIs 1.2 and 1.3, implement pool-scoped xcast modes, and require allow-promisc for requested all-multicast or unicast-promiscuous modes. The VF mailbox can carry only 30 multicast hashes. When ixv has a larger list, request the API 1.2 all-multicast xcast mode instead of extending the legacy SET_MULTICAST message. The PF grants that fallback only to VFs configured with allow-promisc; otherwise ixv reports that only the first 30 addresses are active. Reset xcast state with the VF and have ixv replay the mode implied by its interface flags after multicast updates. Follow DPDK's ixgbe API 1.2/1.3 xcast contract, with allow-promisc policy adapted from igb(4) in a2ed165f0049. MFC after: 1 week Relnotes: yes
The PF advertises the legacy SET_MACVLAN mailbox request but always rejects it. The request installs secondary unicast addresses. Allocate an owned RAR pool for VF secondary addresses, reserve low entries for PF filters, and place VF-primary addresses at the top of the usable RAR range. Reject address collisions and cap each VF at three secondary filters so one guest cannot exhaust the shared table. Clear secondary filters on VF or PF reset and on SR-IOV teardown. This hardware can anti-spoof only the VF primary source address. Reject secondary filters while MAC anti-spoofing is configured, so installing them requires an explicit administrative policy choice. Report optional filter-table allocation failure without disabling SR-IOV. Adapt the owned-RAR allocation and reset-cleanup model from igb(4) in a2ed165f0049 to DPDK's ixgbe SET_MACVLAN mailbox semantics. MFC after: 1 week Relnotes: yes
The shared X550 code provides malicious-driver detection, event decoding, and per-pool recovery operations, but the PF never enables or services them. A malformed VF descriptor can therefore go undetected and avoid the per-pool recovery path supplied by the MAC. Configure IOV state while VF DMA remains disabled, then enable MDD and activate the VFs only after PF queue initialization is complete. On an MDD event, withdraw mailbox CTS and gate the VF pool through PFVFTE and PFVFRE. Retain the per-queue WQBR blocks until the VF enters a new reset epoch; PFVFTE can still permit descriptor fetches into the internal queue, so releasing WQBR early would allow a hostile VF to retrigger MDD before it resets. Send the non-CTS reset notification after servicing the VF mailbox. Let a posted VF request win mailbox arbitration, defer notification if the pass produced a response, and retry failed notifications from the periodic admin pass. Poll WQBR so recovery does not depend on another mailbox interrupt edge, while suppressing already-fenced pools. Latch a PF reset request until the next hardware initialization. The X550 datasheet defines every bit of WQBR_RX and WQBR_TX as a queue bit, so an all-ones value is valid. Reject it only when IXGBE_STATUS, which has reserved-zero bits, also reads as all ones and confirms dead MMIO. Temporarily disable MDD around live multiqueue SRRCTL drop-mode updates, which hardware otherwise reports as queue-context changes. Serialize that window with the iflib context lock and resample pending work after MDD is restored. Apply the per-pool recovery model used by igb(4) in a2ed165f0049 to the existing DPDK-derived X550 hooks. The same register interface is documented for X552 and X553, so cover the entire X550 family. Document that VF traffic remains disabled until the reset handshake completes. MFC after: 2 weeks Relnotes: yes
iflib counts resets initiated by its transmit watchdog in 69c3e0de01c1. Export the counter in the per-device iflib sysctl tree so every driver provides the diagnostic without a driver callback or duplicate storage. A watchdog reset does not establish how many packets failed. It can recover a hardware stall involving several queued packets or a missed completion involving no packet loss. Stop adding one output error per watchdog event in em(4), igb(4), and igc(4). Remove the redundant driver counters and move the diagnostic to dev.<driver>.<unit>.iflib.tx_watchdog_events. MFC after: 1 month Relnotes: yes
- Adds SR-IOV VF status to the existing ifconfig "-v" output - Adds ioctl command for reporting VF status info from drivers - Adds support to iflib for drivers to handle this new ioctl - Add support for ioctl in ixl(4) Signed-off-by: Eric Joyner <erj@freebsd.org> Relnotes: yes Differential Revision: https://reviews.freebsd.org/D19647
Add -L to query the generic packed-nvlist IOV_GET_STATUS interface. Report PF enable state and configured and total VF counts. For each VF, print its PCI address, newbus attachment, bound driver, and ppt state. Retry size negotiation if the topology changes between ioctls and reject malformed or incompatible status records. Keep NIC-specific operational state in ifconfig -v; iovctl owns the device-neutral PCI topology and applies to any SR-IOV device class. Relnotes: yes
Add access and trunk VLAN policy to the SR-IOV schema. Access VFs use a hardware PVID and cannot alter their VLAN membership. Trunk VFs may register up to 16 VLANs, while VLAN 0 remains implicitly admitted for untagged and priority-tagged traffic. Enable hardware VLAN anti-spoofing and maintain the MAC-by-VLAN filter cross-product used by DPDK. Apply Linux's untrusted-VF limits of 18 MAC addresses and 16 VLANs so one guest cannot consume the shared PF filter table without bound. Report the effective policy through the VF status interface and document the iovctl schema. MFC after: 2 weeks Relnotes: yes
10G-BX optics use paired wavelengths to carry 10 Gb/s Ethernet over a single strand of single-mode fiber. Their 10G compliance byte is empty, so identify them from the SFF-8472 nominal signaling rate and single-mode reach fields. When an EEPROM also advertises 1G BASE-BX10, give the complete 10G bitrate and reach signature precedence. Otherwise retain FreeBSD's permissive 1G-BX identification rather than requiring a nominal 1.3 GBd rate. MFC after: 2 weeks Relnotes: yes
The bhyve_config(5) variable `virtio_msix` is namescoped to `virtio.msix`. Configurations that have the old variable will automatically be mapped to the new one, with a warning message printed out. Relnotes: yes Reviewed by: ziaee, markj Differential Revision: https://reviews.freebsd.org/D58390
Reviewed by: manu, adrian Differential Revision: https://reviews.freebsd.org/D58798
PR: https://bugs.freebsd.org/bugzilla/show_bug.cgi?id=296234(exp-run) Relnotes: yes Sponsored by: The FreeBSD Foundation Differential Revision: https://reviews.freebsd.org/D57772
E610 inherits the X550-family virtualization registers, anti-spoofing controls, and malicious-driver operations, but the frontend does not advertise SR-IOV and cannot negotiate the mailbox revision needed by E610 VFs. Initialize the X550-family PF/VF mailbox registers for E610 and use PFVFLREC for its VF reset events, following DPDK shared ixgbe code. Advertise the E610 SR-IOV capability, accept API 1.6 only on E610, carry the existing xcast and queue operations forward to that revision, and return the cached physical link speed and state with the three-dword E610 operation. Unsupported RSS and optional feature requests continue to receive explicit failures. SR-IOV activation also enables the existing X550-derived per-pool MDD recovery path on E610. Document the expanded protection and link-state coverage. Hardware validation created 63 VFs and rejected a 64th without flapping the running PF. Invalid TX and RX descriptor DMA independently asserted the offender's WQBR bit, gated only that VF, preserved sibling traffic, and recovered after the VF reset. FreeBSD ixv, FreeBSD DPDK, Linux ixgbevf, and Linux DPDK exercised the PF mailbox and data paths. MFC after: 2 weeks Relnotes: yes Sponsored by: Dirk-Willem van Gulik from Web Weaving (E610 hardware) Sponsored by: BBOX.io
Add fts_openat() as a new entry point for fts(3). When dirfd is AT_FDCWD the behaviour is identical to fts_open(). Passing a pre-opened directory fd allows fts traversal inside Capsicum capability mode where path-based operations are not permitted. Capability mode users should use fts_parent->fts_dirfd + fts_name with openat(2) to access files. Reviewed by: asomers Relnotes: yes Sponsored by: Google LLC (GSoC 2026) Pull Request: https://github.com/freebsd/freebsd-src/pull/2273
This is ABI-breaking change that could be considered as the bug fix. Requested by: David Timber <dxdt@dev.snart.me> Reviewed by: emaste, markj Sponsored by: The FreeBSD Foundation Relnotes: yes MFC after: 1 week Differential revision: https://reviews.freebsd.org/D58942
Run the file-hierarchy traversal inside Capsicum capability mode using the fts_openat(3) API. This confines chflags to the directory hierarchies named on the command line, so a malicious or buggy tree cannot redirect it at files elsewhere via a crafted symbolic link. Because AT_FDCWD is rejected in capability mode, a directory descriptor for the parent directory of every argument with an absolute path or a path containing ".." is opened once before cap_enter(). Once every descriptor is open, cap_enter() is called and the traversal acts through fd-relative operations: chflagsat(fts_parent->fts_dirfd, fts_name). With -L chflags follows symbolic links, which may point outside the named hierarchy; chflag now rejects such accesses. The new --dereference-links-unsafely option disables the sandbox to restore the historical behavior for the rare callers that rely on it. But if the symbolic was link was named directly on the command line, chflags will still follow it (unless -h was given). Add functional tests covering relative, absolute, "..", recursive and mixed path arguments, and the symlink handling in both the default and --dereference-links-unsafely modes; they skip on filesystems that do not support the uchg flag. Sponsored by: Google LLC (GSoC 2026) Reviewed by: asomers Relnotes: yes (for the -L behavior change) Pull Request: https://github.com/freebsd/freebsd-src/pull/2375
In VM and cloud environments it is often possible to enlarge virtual
disks; this can be useful, for example, if a system is launched with a
small root disk and it later becomes clear that more space is needed.
On kernels which support run-time resizing of disks (for NVMe, this was
added in November 2025; some other disk types have supported this for
longer) a SIZECHANGE notification is sent to userland via devd.
Add a "nostart" rc.d script (runnable manually but not automatically at
boot time) and a devd script which invokes it when a notification
arrives. The rc.d script enlarges the "final partition" on partitioned
geoms, or the UFS filesystem or zpool device when triggered on a disk
containing either of those.
Reviewed by: imp, ziaee
MFC after: 2 weeks
Relnotes: Disk partitions and filesystems can be enlarged
automatically when disks grow by setting
growfs_postboot_enable=YES in /etc/rc.conf.
Sponsored by: Amazon
Differential Revision: https://reviews.freebsd.org/D58582
Enable the new growfs_postboot mechanism. Note that this also implies disabling automatic allocation of swap space on the root disk, since we cannot grow the root filesystem if swap space is allocated after it. This will not be MFCed since it is a significant behavioural change. Sponsored by: Amazon Relnotes: yes
Manually specified eui64 value gets converted to big endian twice: first using htobe64() and then using be64enc(). On little-endian hosts that results in a little-endian value instead of a big-endian. Fix by removing htobe64() for a user submitted value. Fixes: https://cgit.freebsd.org/src/commit/?id=409a80e5a434 ("bhyve: Create EUI64 for NVMe namespaces") Reviewed by: chuck Relnotes: yes Sponsored by: The FreeBSD Foundation MFC after: 3 weeks Differential Revision: https://reviews.freebsd.org/D59080
In the world of containers, mounting a unix(4) socket is a common practice to allow communication between processes within containers. For example, both Podman and Docker can expose a unix(4) socket, and that same unix(4) socket can be mounted as a file accessible to a process inside a container, allowing that application to control Podman or Docker. Another example is PHP-FPM with NGINX, where, instead of using TCP/IP for communication between containers, a unix(4) socket is sufficient. However, nullfs(4) and all related components do not allow mounting a VSOCK on top of another. The current workaround involves creating the socket in a directory and mounting that directory. This is an option, though it does not provide a good user experience compared to directly mounting a VSOCK on top of another, since the application that creates the socket may create other sockets in that directory, and the user may not wish to share them, or, worse yet, applications that create unix(4) sockets may not provide any authentication at all, as they may assume that security at the file system level is sufficient. Reviewed by: dfr@ Approved by: dfr@ Relnotes: yes Differential Revision: https://reviews.freebsd.org/D59158
Pushed using the RTL8723BU. Reviewed by: ziaee, avos, adrian Relnotes: yes Differential Revision: https://reviews.freebsd.org/D59205
The default on my laptop is annoyingly bright, and this is a useful feature to mitigate that. The backlight script is largely a copy of the mixer service which provides the same value for mixers, but this one is specifically dependant on kld to allow DRM drivers a chance to attach. Note that it's off by default to avoid interference with DEs, and document the capability in backlight(8). Set backlight_enable=YES in rc.conf(5) to enable save/restore. Relnotes: maybe Reviewed by: bapt, ivy, manu, ziaee Differential Revision: https://reviews.freebsd.org/D59296
Make the witness LOCK_CHILDCOUNT a configurable kernel option. On machines with a very high core count the default value is too low, leading to witness exhaustion after boot. Relnotes: yes Reviewed by: kib, ziaee Signed-off-by: Kajetan Puchalski <kajetan.puchalski@arm.com> Closes: https://github.com/freebsd/freebsd-src/pull/2398
Add iri_rcv_tstmp to if_rxd_info so an isc_rxd_pkt_get() driver can report a hardware RX timestamp. Copy it into m_pkthdr.rcv_tstmp, reusing the generic mbuf timestamp path. Widen iri_flags from uint8_t to uint32_t and define the flags drivers may supply. Mask the flags before copying them into the mbuf so no other mbuf state can leak through the driver callback. Place the timestamp next to iri_frags to avoid an alignment hole, and document its nanoseconds-since-boot representation and validity flags. Bump __FreeBSD_version because changing if_rxd_info breaks KBI. Reviewed by: gallatin Signed-off-by: Sreekanth Reddy <sreekanth.reddy@broadcom.com> Differential Revision: https://reviews.freebsd.org/D58638
Add NIC-specific VF status to the existing ifconfig -v output. Fetch the data through libifconfig using a separate native route Netlink query. Group optional identity, initialization, resources, VLAN policy, administrator policy, protocol, traffic-permission, and fault containment fields. Omitted fields remain distinct from false or zero. Refer users to iovctl -L for device-neutral PCI attachment and passthrough state. This is a Netlink-native evolution of the original interface by Eric Joyner. Relnotes: yes Sponsored by: Intel Corporation (initial version) Sponsored by: BBOX.io Differential Revision: https://reviews.freebsd.org/D58778 Co-authored-by: Eric Joyner <erj@FreeBSD.org>
Display top-like NFS server I/O using dtrace(1). Also supports JSON output for time-series. Relnotes: yes Reviewed by: ziaee, bcr, adrian Differential Revision: https://reviews.freebsd.org/D59438
Desktop AMIs have xrdp enabled and boot to a KDE desktop; they are as compatible as possible with EC2 Windows AMIs, setting a random password and printing it to the console in encrypted format to be retrieved using the EC2 GetPasswordData API. Two rc.d scripts are included in this commit which will not exist in the long term: ec2_addpass will become part of the ec2-scripts package, and ec2_desktop_extras will go away once its functionality is included elsewhere. MFC After: 1 month Relnotes: yes Sponsored by: Amazon
Accept and validate the ipv6only option. When dhclient receives this option and IPv6 connectivity is available, stop the DHCP configuration process and wait for the duration specified by the option before restarting DHCP discovery. If the address was previously leased, disassociate it and send a DHCPRELEASE packet. Use netlink to check for IPv6 connectivity. Also, unregister ignored options from default PRL. Reviewed by: ziaee, kfv Tested by: Marek Zarychta <zarychtam@plan-b.pwste.edu.pl> Relnotes: yes Differential Revision: https://reviews.freebsd.org/D56637
tpm(4) was removed from amd64 GENERIC because it broke suspend and resume. The preceding lifecycle, state-save, interrupt, locality, and teardown fixes address those failures for both TPM 1.2 and TPM 2.0. Restore the driver to amd64 GENERIC and MINIMAL, where TPM entropy harvesting remained enabled. Enable the driver and entropy harvesting in the MPC85XX and QORIQ64 configurations, which already provide FDT, spibus, and the platform SPI controller required by FDT-attached TPMs. Leave the generic AIM and POWER configurations unchanged because they have no TPM attachment bus. The TPM 1.2 path completed repeated S3 cycles and command tests on ThinkPad T430 and T440p systems. The TPM 2.0 path completed repeated device and full-system suspend/resume cycles on a ThinkPad P51. The PowerPC configuration matrix was checked to retain tpm(4) only where its FDT SPI attachment path is present. PR: https://bugs.freebsd.org/bugzilla/show_bug.cgi?id=291067 Reviewed by: kevans Tested by: Marek Zarychta <zarychtam_plan-b.pwste.edu.pl> (tpm1.2) Fixes: https://cgit.freebsd.org/src/commit/?id=16f8ea6a81b5 ("amd64: Remove tpm(4) from GENERIC for now") MFC after: 1 month Relnotes: yes Sponsored by: BBOX.io Differential Revision: https://reviews.freebsd.org/D59247
MFC After: 1 week Approved by: dteske Reviewed by: dteske Differential Revision: https://reviews.freebsd.org/D59658
The size of multiple embedded structs have changed and may lead to problems (pci_error_handlers in pci_driver, dev_pm_info in struct device). Allow these changes to be detected by bumping __FreeBSD_version. MFC after: 3 days
mfi_drive_name() built the "Exx:Syy" drive location string using struct mfi_pd_info's encl_index field, the enclosure's firmware- internal position index. Broadcom's own storcli/MegaCli tooling instead leads with the enclosure's Device ID (EID) in its primary drive listing; encl_index only shows up as "Position" in a detailed per-enclosure view. Both numbers are raw, unmodified firmware values already fetched into struct mfi_pd_info/mfi_pd_address, but only encl_index was ever displayed or accepted as input, leading to confusion when cross-referencing drive locations against storcli output. Switch mfi_drive_name() and mfi_lookup_drive() to use encl_device_id instead, aligning FreeBSD's enclosure numbering with Broadcom's own utilities. Since mrsasutil(8) is the same binary as mfiutil(8) under a different name, this applies to both mfi(4) and mrsas(4) alike. This is a user-visible behavior change: the numeric value of "xx" in "Exx:Syy" now differs from before for any enclosure whose EID and position index don't match, affecting anyone scripting against the previous numbering. PR: https://bugs.freebsd.org/bugzilla/show_bug.cgi?id=294353 Reviewed by: imp Relnotes: yes Differential Revision: https://reviews.freebsd.org/D59654
Use the Hyper-V reset/MAC exchange for 82576 and I350 VFs instead of the native posted mailbox protocol, which the Windows PF does not service. Read the host assigned address through configuration bytes 0x201 through 0x206 only during reset, and use it to identify the matching synthetic hn(4) interface. The operations are local to the VF frontend. Poll hardware link status rather than retaining a native mailbox link handshake. Leave MAC, multicast, promiscuous-mode, and VLAN membership policy with the host. Disable guest VLAN registration and native receive limit requests, and limit the VF to an MTU of 1500 bytes. Preserve accumulated statistics across host resets without counting a counter clear as a wrap. Reject inaccessible register samples and rebase after a reset indication or a disabled transmit queue, including when the PF blocks the queue for malicious driver detection. Document single queue support and host assigned access VLANs. Guest VLAN trunks are not supported: tagged transmissions with the Windows I350 PF driver 14.1.5.0 can disable VF queues even on a trunk-configured port, and the limitation was also reproduced with Windows VF drivers. Such trunks must use the synthetic path with SR-IOV disabled for that virtual adapter. Tested on Windows Server 2025 Hyper-V with both 82576 and I350 PFs. Relnotes: yes Sponsored by: BBOX.io
The Unicode closing single quotation mark is classified as a homoglyph and can trip automated code quality checks in downstream CI pipelines or cause code review UIs to refuse to display a file. If used as an apostrophe, use the ASCII single quote instead. If used as a closing single quote, replace with double quotes or no quotes at all. Sponsored by: Klara, Inc. Sponsored by: NetApp, Inc. Reviewed by: ziaee, obiwac, olce Differential Revision: https://reviews.freebsd.org/D59911
Commits about commands found in man section 1 (other than networking).
This internet draft (which is close to being an RFC) specifies a new NFSv4.2 attribute which tells the NFSv4.2 client to not cache file data. (Similar to O_DIRECT, but triggered by this attribute set on the file on the NFSv4.2 server and not by the application's open(2).) https://datatracker.ietf.org/doc/draft-ietf-nfsv4-uncacheable-files/ This patch adds a new chflags(1) flag called UF_DONTCACHE to implement this. Patches for NFS and ZFS will be done separately. Reviewed by: kib MFC after: 2 weeks Differential Revision: https://reviews.freebsd.org/D58181
Unlike its GNU counterpart, our tail(1) has always errored out if given repetitive or contradictory options, even prior to Keith Bostic's 1991 reimplementation. There is no good reason to continue to do so, not even tradition, since many other commands (including head(1)) simply apply the rightmost option in cases like this. MFC after: 1 week Reviewed by: allanjude, markj Differential Revision: https://reviews.freebsd.org/D58192
Now that fetchTimeout works reliably, setting an alarm is not only no longer necessary but counterproductive, as it will trigger even if the connection is not actually stalled but merely slow. While here, improve the wording of the manual page's description of the various options for setting a timeout. MFC after: 1 week Reviewed by: op Differential Revision: https://reviews.freebsd.org/D57911
While preparing GPT-schemed RaspberryPi images for the NanoBSD Reimagined GSoC 2026 project, a discrepancy was identified between mkimg(1) and gpart(8) regarding Microsoft Basic Data partitions (GUID !ebd0a0a2-b9e5-4433-87c0-68b6b72699c7). Currently, mkimg(1) relies on the MBR-centric name "ntfs" to identify this partition type under the GPT scheme. Conversely, gpart(8) identifies this type as "ms-basic-data". To allow automation scripts (such as those consuming from gpart backup) to use a common partition type across tools, add ALIAS_MS_BASIC_DATA as a valid alias. This is part of a larger effort to avoid a custom, MBR-based image generation logic for embedded SoCs like the Raspberry Pi, standardizing on GPT layouts across all supported FreeBSD embedded devices. Reviewed by: imp MFC after: 2 weeks Differential Revision: https://reviews.freebsd.org/D58198
This allows to use output of '/usr/bin/time -ao foo' as direct input to ministat(1). While here make diagnostic message more verbose.
GNU hexdump supports octal and hex, we add supports for BSD style hexdump for better compatibility. See: https://github.com/llvm/llvm-project/pull/206581/ MFC after: 2 weeks Sponsored by: The FreeBSD Foundation Differential Revision: https://reviews.freebsd.org/D58074
- gnum4.c: fix m4_warnx() to use vwarnx() instead of warnx() - eval.c: improve error messages for empty macro names - extern.h: remove compute_prevep() declaration - Update OpenBSD version strings MFC After: 3 days
This partially reverts commit 77a201b1705dbd97ea9ebe5b25b1d4ddac8a7d38. Requested by: des, fuz
The fallback glyph is stored at index 0, and does not need to be inserted into a mapping. Previously there was a dead store of add_glyph's return value for the fallback case, which upset Clang's static analyzer. Now, cast the return value to (void) to make it clear this is intentional. Also change add_glyph's fallback parameter to a c99 bool to make its use more clear. Sponsored by: The FreeBSD Foundation Differential Revision: https://reviews.freebsd.org/D57174
install: Allow installing stdin If from_name is "/dev/stdin" or "-" and the target is not a directory, skip the comparison and copy data from standard input to the target. MFC after: 1 week Reviewed by: imp Differential Revision: https://reviews.freebsd.org/D58348
install: Fix typo MFC after: 1 week Reported by: markj Fixes: https://cgit.freebsd.org/src/commit/?id=d34870708db9 ("install: Allow installing stdin")
This is mainy focused on using bool for booleans but also renames some variables for clarity, adds some explicit comparisons, adds some braces, with miscellanous style fixes thrown in. MFC after: 1 week Reviewed by: imp Differential Revision: https://reviews.freebsd.org/D58355
Check the `fdopen` return value before calling `cook_cat`. Reviewed by: markj, bnovkov Differential Revision: https://reviews.freebsd.org/D57741 MFC after: 1 week
pwait: Optionally wait until process is reaped If the new -r option is specified, wait until the target process not only terminates but is reaped. MFC after: 1 week Sponsored by: Klara, Inc. Sponsored by: NetApp, Inc. Reviewed by: kib, markj Differential Revision: https://reviews.freebsd.org/D58314
pwait: Add a SIGINFO handler On SIGINFO, print a space-separated list or remaining processes to standard error. MFC after: 1 week Sponsored by: Klara, Inc. Sponsored by: NetApp, Inc. Reviewed by: kib, markj Differential Revision: https://reviews.freebsd.org/D58386
rc.subr: Fix premature return from wait_for_pids Use pwait's new -r option to wait until the target processes have not only terminated, but also been reaped. PR: https://bugs.freebsd.org/bugzilla/show_bug.cgi?id=293183 MFC after: 1 week Sponsored by: Klara, Inc. Sponsored by: NetApp, Inc. Reviewed by: markj Differential Revision: https://reviews.freebsd.org/D58391
Bump dates Fixes: https://cgit.freebsd.org/src/commit/?id=c8f5e6819d4d ("pwait: Optionally wait until process is reaped") Fixes: https://cgit.freebsd.org/src/commit/?id=eddd8aa99ca8 ("pwait: Add a SIGINFO handler") Fixes: https://cgit.freebsd.org/src/commit/?id=356d0b79cf6f ("rc.subr: Fix premature return from wait_for_pids")
Sponsored by: AFRL, DARPA
Several functions were using sprintf() to write RPC server-controlled data to a stack buffer. Adopt some minimal changes from NetBSD to avoid the potential overflows. Security: CVE-2026-16277 Security: CVE-2026-16461 Reviewed by: khorben MFC after: 1 week Sponsored by: The FreeBSD Foundation Differential Revision: https://reviews.freebsd.org/D58441
Reviewed by: emaste MFC after: 1 week Sponsored by: The FreeBSD Foundation Differential Revision: https://reviews.freebsd.org/D58442
On some platforms, e.g. Linux Clang 22.1.8 / glibc 2.43, strchr() now implements the C23 behaviour where passing a const pointer to strchr() also returns a const pointer. This breaks rpcgen during the bootstrap build, since it assumes the return value is always a mutable pointer. For mkfile_output(), the pointed-to value is never modified, so fix this by making the pointer const as well. For open_log_file(), the current code modifies the supposedly const value in-place to remove the filename suffix, which happens to work but is wrong even in older versions of C. Change the code to use a printf "%.*s" format specifier to strip the suffix instead. MFC after: 1 week Reviewed by: brooks Sponsored by: The FreeBSD Foundation Differential Revision: https://reviews.freebsd.org/D58489
On some platforms, e.g. Linux Clang 22.1.8 / glibc 2.43, strchr() now implements the C23 behaviour where passing a const pointer to strchr() also returns a const pointer. This breaks sort during the bootstrap build, since it assumes the return value is always a mutable pointer. As the returned pointer is never used to modify the value, fix this by making the temporary variable const. MFC after: 1 week Reviewed by: markj Sponsored by: The FreeBSD Foundation Differential Revision: https://reviews.freebsd.org/D58491
On some platforms, e.g. Linux Clang 22.1.8 / glibc 2.43, strchr() now implements the C23 behaviour where passing a const pointer to strchr() also returns a const pointer. This breaks xinstall during the bootstrap build, since it assumes the return value is always a mutable pointer. As the returned pointer is never used to modify the value, fix this by making the temporary variable const. MFC after: 1 week Reviewed by: ray, markj, emaste Sponsored by: The FreeBSD Foundation Differential Revision: https://reviews.freebsd.org/D58492
On some platforms, e.g. Linux Clang 22.1.8 / glibc 2.43, strchr() now implements the C23 behaviour where passing a const pointer to strchr() also returns a const pointer. This breaks mkimg during the bootstrap build, since it assumes the return value is always a mutable pointer. Make the existing 'sep' pointer const to fix the first case, and for the second, introduce a new non-const pointer for strchr, since we do modify the result in that case. MFC after: 1 week Reviewed by: markj Sponsored by: The FreeBSD Foundation Differential Revision: https://reviews.freebsd.org/D58493
On some platforms, e.g. Linux Clang 22.1.8 / glibc 2.43, strchr() now implements the C23 behaviour where passing a const pointer to strchr() also returns a const pointer. This breaks m4 during the bootstrap build, since it assumes the return value is always a mutable pointer. Since the returned value is never modified, simply make the temporary const. MFC after: 1 week Reviewed by: bapt, dim Sponsored by: The FreeBSD Foundation Differential Revision: https://reviews.freebsd.org/D58494
Add some minimal handling of category sources other than static kernel sources. We don't actually look up dynamic sources yet (that would require extended trace records to add the file names to the trace file since we can't assume the trace file is running on a kernel with the same numbers.) Make the decision to append a "src/" prefix to each file name dependent on the category source. Reviewed by: kib Sponsored by: Innovate UK Differential Revision: https://reviews.freebsd.org/D58412
* On SIGINFO, print the current path to stderr rather than stdout. * Do so immediately, instead of the next time we finish a directory. * Document this behavior in the manual page. PR: https://bugs.freebsd.org/bugzilla/show_bug.cgi?id=296861 MFC after: 1 week Fixes: https://cgit.freebsd.org/src/commit/?id=d1588599c024 ("Report the next directory being scanned ...") Reviewed by: wollman Differential Revision: https://reviews.freebsd.org/D58702
Reviewed by: fuz, ngie Co-authored-by: Robert Clausecker <fuz@FreeBSD.org> Differential Revision: https://reviews.freebsd.org/D58727
The current default has been unchanged for 14 years. Increase it to keep pace with modern hardware and software. security/pinentry-gnome, in particular, can sometimes need 112 kB. PR: https://bugs.freebsd.org/bugzilla/show_bug.cgi?id=297452 MFC after: 2 weeks Sponsored by: ConnectWise Reviewed by: cye, emaste Differential Revision: https://reviews.freebsd.org/D58811
Our buffer is half a megabyte, but we are only initializing the first two bytes. Switching from static to dynamic initialization moves it from .data to .bss, greatly reducing the size of the binary. Fixes: https://cgit.freebsd.org/src/commit/?id=cf74b63d61b4 ("yes: Completely overengineer") MFC after: 1 week Sponsored by: Klara, Inc. Reviewed by: kevans Differential Revision: https://reviews.freebsd.org/D58890
Run the file-hierarchy traversal inside Capsicum capability mode using the fts_openat(3) API. This confines chflags to the directory hierarchies named on the command line, so a malicious or buggy tree cannot redirect it at files elsewhere via a crafted symbolic link. Because AT_FDCWD is rejected in capability mode, a directory descriptor for the parent directory of every argument with an absolute path or a path containing ".." is opened once before cap_enter(). Once every descriptor is open, cap_enter() is called and the traversal acts through fd-relative operations: chflagsat(fts_parent->fts_dirfd, fts_name). With -L chflags follows symbolic links, which may point outside the named hierarchy; chflag now rejects such accesses. The new --dereference-links-unsafely option disables the sandbox to restore the historical behavior for the rare callers that rely on it. But if the symbolic was link was named directly on the command line, chflags will still follow it (unless -h was given). Add functional tests covering relative, absolute, "..", recursive and mixed path arguments, and the symlink handling in both the default and --dereference-links-unsafely modes; they skip on filesystems that do not support the uchg flag. Sponsored by: Google LLC (GSoC 2026) Reviewed by: asomers Relnotes: yes (for the -L behavior change) Pull Request: https://github.com/freebsd/freebsd-src/pull/2375
Fix a documentation discrepancy, where the implementation was updated to use uncertainty propagation for the ratio of means, but the example output in the manual page was left unchanged. Update the manual page example from 70.7384% to 102.3% to reflect the actual output. While here, also update the example in the README. Reviewed by: ziaee Fixes: https://cgit.freebsd.org/src/commit/?id=a304ad90e9ae ("Reduce the bogosity of ministat's % difference calculations.") MFC after: 1 week Differential Revision: https://reviews.freebsd.org/D59157
* Fix case where the source is - and the target exists. * Only call chflags() (to remove flags that might prevent us from replacing an existing target) in the exists case; otherwise, to_sb.st_flags is uninitialized. * Rename the source file in the stdin test case. * Extend null and stdin test cases to cover the case where the target already exists. PR: https://bugs.freebsd.org/bugzilla/show_bug.cgi?id=297681 MFC after: 1 week Fixes: https://cgit.freebsd.org/src/commit/?id=d34870708db9 ("install: Allow installing stdin") Reviewed by: markj Differential Revision: https://reviews.freebsd.org/D59144
This internet draft (which is close to being an RFC) specifies a new NFSv4.2 attribute which tells the NFSv4.2 client to not cache file data. (Similar to O_DIRECT, but triggered by this attribute set on the file on the NFSv4.2 server and not by the application's open(2).) https://datatracker.ietf.org/doc/draft-ietf-nfsv4-uncacheable-files/ This patch adds a new chflags(1) flag called UF_NOCACHE to implement this. Patches for NFS and ZFS will be done separately. This is a redo of the patch, with a requested name change and a #ifdef in strtofflags.c so that it doesn't break some Linux cross build. The name change was requested by fuz@. Reviewed by: kib (earlier version) MFC after: 2 weeks Differential Revision: https://reviews.freebsd.org/D58181
With no file argument, fortune looks for a database named fortunes in FORTDIR. The base system has not shipped that file since 0538d7bbe620 (FreeBSD 12), only freebsd-tips, so the default invocation failed even though a valid database remained. Callers such as xlockmore's marquee and nose modes (fortune -s) then displayed the error as the epigram. If the named fortunes file is absent, scan every database in the existing search path. /usr/local/share/games/fortune stays on that path so fortune-mod-* packages keep working; when fortune-mod-freebsd-classic restores the fortunes file, it is still preferred. fortune -f with no arguments lists the same files that would be searched. MFC after: 1 week Reviewed by: ziaee, fuz Differential Revision: https://reviews.freebsd.org/D59057
Support cat -A, -E and -T, which are commonly used by Linux shell scripts. -E prints a "$" at the end of each line, -T renders tabs as ^I, and -A is equivalent to -vET. MFC After: 1 week Discussed with: jrtc27 Reviewed by: jrtc27, ziaee Differential Revision: https://reviews.freebsd.org/D59250
Prior to FreeBSD 11, vm_cnt was named cnt. When it was renamed, vmstat was modified to fall back to the old name if the new name was not found. It's time we dropped this. Reviewed by: kib, jhb, emaste Differential Revision: https://reviews.freebsd.org/D59255
truss reports every system call a process makes, which for anything
larger than a toy program buries the calls of interest. Add -t, taking
a comma-separated expression naming the system calls to report.
A term is the name of a system call, which may contain the fnmatch(3)
wildcards; a system call number in decimal; or "@group" naming a group
of related system calls. A term prefixed with '!' excludes what that
one term matches rather than including it, and applies to no other
term. An expression whose terms are all negated subtracts from the set
of every system call; any other expression selects from an empty one.
Terms apply in order and the last one to match a system call decides
whether it is reported. Repeating -t appends, so "-t a -t b" and
"-t a,b" are equivalent. An empty term is ignored, so an empty
expression filters nothing and a stray comma is not an error.
truss -t @file,@net fetch https://www.freebsd.org/
truss -t '!@memory' make buildworld
truss -t '@desc,!@read,!@write' -p 34
truss -c -t 'readlink*' /bin/ls
truss -t 3,4 cat file
A number selects the system call with that number in the ABI of the
traced process, so "-t 3" selects read(2) from a native process but
close() from a Linux one.
Thirteen groups are provided to start with: @all, @creds, @desc, @file,
@ipc, @memory, @net, @none, @proc, @read, @signal, @time and @write.
These were derived by working through sys/kern/syscalls.master; the
audit event in that file is too sparse to drive the grouping by itself
(316 distinct events over 509 live system calls, 111 of them AUE_NULL),
so the groups are curated, but they are curated as patterns rather than
as name lists. A family sharing a naming convention is written as one
pattern -- "extattr_*_file", "__acl_*_fd", "sctp_*" -- so system calls
added later join the right group without further change here.
A group's member list is an expression in exactly the form -t accepts,
so a group can say anything a user can say on the command line: a
member may be a pattern, a number or another group, and may be negated.
@desc is built from @read and @write without repeating them, and @none
is the single member "!*". Keeping the two languages identical means a
group defined from a -t expression supplied elsewhere needs no
translation to become a member list. Adding a group is a member list
plus one entry in syscall_groups[].
"truss -t" with no expression prints the groups and exits.
Matching is done against the name truss displays and against that name
with any compatibility or ABI prefix removed, so @file selects
compat11.stat, freebsd32_stat and linux_newstat as well as stat.
A name or pattern matching no system call of any ABI truss understands
is reported with a warning, since it is almost always a typo, but it is
kept and simply never matches. sysdecode(3) is the oracle: it names
every system call of every such ABI whether or not that ABI's module is
loaded, and names them exactly as truss reports them. Numbers are not
checked this way. A process may issue any number the kernel can hold,
whether or not a system call is implemented behind it; one that is not
returns ENOSYS, which truss reports like any other result. Only a
number too large to be one at all is rejected.
A system call excluded by -t is not decoded, so the filter also removes
the cost of formatting arguments that would never be printed, and it is
left out of the -c summary.
The option letter is the one truss on System V Release 4 and SunOS uses
for this feature, "-t [!]syscall,...", and which truss(1) already names
as its model. The syntax is deliberately not bug-compatible with it:
there '!' is sticky for the remainder of a list, and a second -t
discards the first when the first began with '!'.
usr.bin/truss/tests is new, so etc/mtree/BSD.tests.dist gains an entry.
Without -t the behaviour is unchanged.
MFC after: 2 weeks
Reviewed by: fuz
Differential Revision: https://reviews.freebsd.org/D59275
Commit 50c1240ebfaf moved the offset parsing into the PART_KIND_FILE case of
the switch, leaving PART_KIND_SIZE with no offset handling.
The offset was then silently ignored, so "-p efi::$size:$start" as used by
release/${ARCH}/mkisoimages.sh packed the ESP immediately after the preceding
partition.
Parse the offset outside the switch so both forms honour it.
Add tests covering absolute and relative offsets in both forms.
Reviewed by: jrtc27, bsdimp, jlduran
Approved by: jlduran, bsdimp
Sponsored by: Netflix
Assisted-by: Claude Code (Opus 5)
Reviewed by: ziaee Event: Berlin Hackathon 202609
Reviewed by: fuz, oshogbo Approved by: fuz (mentor) Pull Request: https://github.com/freebsd/freebsd-src/pull/1489
Respect PORTSDIR variable for those who have the ports collection in a different place than /usr/ports. PORTSDIR is a very common variable used in the ports framework and in /etc/make.conf among other places. While here, remove and old reference to CVS. Reviewed by: delphij@, ngie@ Approved by: ngie@ Differential Revision: https://reviews.freebsd.org/D42156
Reviewed by: fuz, dteske Approved by: fuz (mentor), dteske (mentor) MFC after: 1 week Differential Revision: https://reviews.freebsd.org/D59666
The new ptrace(2) features allow to change truss(1) to systematically operate on the process descriptors instead of pids. Allocate the global kqueue that tracks all noted children by pdopenpid()-ing them and adding to the kqueue with EVFILT_PROCDESC/NOTE_PDSIGCHLD. The activated knote triggers the pdwait() call to return the child tracing info. This replaces the waitid(P_ALL) call in the non-capsicumized truss(1) eventloop. Reviewed by: markj Sponsored by: The FreeBSD Foundation MFC after: 1 week Differential revision: https://reviews.freebsd.org/D58094
Reviewed by: fuz, dteske Approved by: fuz (mentor), dteske (mentor) MFC after: 1 week Differential Revision: https://reviews.freebsd.org/D59754
Reviewed by: fuz Approved by: fuz (mentor) MFC after: 1 week Differential Revision: https://reviews.freebsd.org/D59773
Reviewed by: fuz Approved by: fuz (mentor) MFC after: 1 week Differential Revision: https://reviews.freebsd.org/D59774
MFC after: 3 days Approved by: bnovkov (mentor) Sponsored by: fme AG Differential Revision: https://reviews.freebsd.org/D59927
Simplify the way we build paths. Avoid decolonification of source paths. Remove gnu directories and add non-tracked ones. Fix a memory leak while here. Approved by: ngie@ Differential Revision: https://reviews.freebsd.org/D59846
Man pages, release notes, etc.
While here, remove the long-unused dash in the first line. Reviewed by: ziaee, olce Fixes: https://cgit.freebsd.org/src/commit/?id=ddf144a04b53 ("ps.1: Revamp: Explain general principles, update to match reality") MFC after: 1 day Differential Revision: https://reviews.freebsd.org/D58038
* Modernize the markup * Describe the comment syntax * Drop obsolete advice * Capitalize sentences * Improve the language * Replace no_tld_query with no-tld-query; both are supported, but all the other multi-word options use hyphens rather than underscores. * Add missing ENVIRONMENT section * Redo the example MFC after: 1 week Reviewed by: markj Differential Revision: https://reviews.freebsd.org/D57921
MFC after: 1 week
Reviewed by: markj Sponsored by: The FreeBSD Foundation MFC after: 1 week Differential revision: https://reviews.freebsd.org/D57124
Also do not start a new list for each flag item. Reviewed by: markj MFC after: 1 week Sponsored by: The FreeBSD Foundation Differential revision: https://reviews.freebsd.org/D57124
Submitted by: des MFC after: 1 week
Submitted by: des MFC after: 1 week
Reviewed by: mckusick Discussed with: markj Tested by: pho Sponsored by: The FreeBSD Foundation MFC after: 1 week Differential revision: https://reviews.freebsd.org/D57658
Commit 74654ba3b1b3 added and new chflags(1) flag called "udontcache" or "dontcache". This patch documents this flag. This is a content change. Reviewed by: kib MFC after: 2 weeks Differential Revision: https://reviews.freebsd.org/D58181
Reviewed by: markj Sponsored by: The FreeBSD Foundation MFC after: 1 week Differential revision: https://reviews.freebsd.org/D58123
Document the global fetchTimeout variable, now that it works reliably. MFC after: 1 week Reviewed by: op Differential Revision: https://reviews.freebsd.org/D57910
Reviewed by: markj Sponsored by: The FreeBSD Foundation MFC after: 1 week Differential revision: https://reviews.freebsd.org/D58247
According to RFC 1918, the following IP prefixes are reserved for
private internets:
10.0.0.0/8
172.16.0.0/12
192.168.0.0/16
This PR fixes the prefix lengths in references to private networks
("RFC 1918 networks", "the standard private IP address ranges").
The changes are limited to man pages.
Signed-off-by: Yusuke Ichiki <public@yusuke.pub>
Pull Request: https://github.com/freebsd/freebsd-src/pull/2328
MFC after: 3 days
Reviewed by: markj Sponsored by: The FreeBSD Foundation MFC after: 1 week Differential revision: https://reviews.freebsd.org/D58292
boot.9 was moved to kern_reboot.9, but this reference was not changed appropriately. Reviewed by: mhorne, kib, emaste Fixes: https://cgit.freebsd.org/src/commit/?id=800e74955d4e ("boot(9): update to match reality") MFC after: 3 days Differential Revision: https://reviews.freebsd.org/D58350
Reviewed by: markj Sponsored by: The FreeBSD Foundation MFC after: 1 week Differential revision: https://reviews.freebsd.org/D58315
Point out which features are non-POSIX and thus can not be safely assumed to be portable and exist in other implementations. Relnotes: YES! Reviewed by: ziaee, jilles Differential Revision: https://reviews.freebsd.org/D55333
List every AQC part aq_vendor_info_array[] probes, each with the maximum speed aq_hw_capabilities() grants it. Only the Atlantic 2 parts link at 10 Megabit. The AQC100 and AQC100S are the only SFP+ controllers; the rest are twisted pair. Reviewed by: adrian, ziaee Signed-off-by: Nick Price <nick@spun.io> Differential Revision: https://reviews.freebsd.org/D58144
Reviewed by: ziaee, imp Differential Revision: https://reviews.freebsd.org/D58267
contigmalloc.9: Note that M_WAITOK may still return NULL Reviewed by: markj, bapt Sponsored by: The FreeBSD Foundation Differential Revision: https://reviews.freebsd.org/D58382
contigmalloc.9: Correct typo Reported by: alc, rlibby Fixes: https://cgit.freebsd.org/src/commit/?id=caabdb3aefdc ("contigmalloc.9: Note that M_WAITOK may still return NULL")
The manual page claimed that SIGINFO caused information to be printed to stdout, when in fact it is printed to stderr, as one would expect. This has been true ever since the feature was first added in 2003. MFC after: 1 week Fixes: https://cgit.freebsd.org/src/commit/?id=00d321a2b395 ("Add a SIGINFO handler.") Reviewed by: jilles Differential Revision: https://reviews.freebsd.org/D58392
Describe the disabled, adaptive, and low-latency settings and their interrupt-rate tradeoffs. MFC after: 1 week
Describe the disabled, adaptive, and low-latency settings and their interrupt-rate tradeoffs. MFC after: 1 week
Reviewed by: markc Sponsored by: The FreeBSD Foundation MFC after: 1 week Differential revision: https://reviews.freebsd.org/D58458
mknod.2: update the man page State that FIFOs can be created, document the requirement that dev must be zero then. Mention whiteouts. PR: https://bugs.freebsd.org/bugzilla/show_bug.cgi?id=297082 Reviewed by: emaste Sponsored by: The FreeBSD Foundation MFC after: 3 days Differential revision: https://reviews.freebsd.org/D58478
mknod.2: properly document root requirements Submitted by: Martijn Dekker <mcdutchie@hotmail.com> PR: https://bugs.freebsd.org/bugzilla/show_bug.cgi?id=297082 Fixes: https://cgit.freebsd.org/src/commit/?id=4090d103b0c3 ("mknod.2: update the man page") MFC after: 3 days
Reviewed by: markj Sponsored by: The FreeBSD Foundation MFC after: 1 week Differential revision: https://reviews.freebsd.org/D58463
Reviewed by: markj Sponsored by: The FreeBSD Foundation MFC after: 1 week Differential revision: https://reviews.freebsd.org/D58463
There are expected to be additional improvements and this RELNOTES entry will be updated accordingly.
Reviewed by: des MFC after: 1 week Differential Revision: https://reviews.freebsd.org/D58483
Adjust the man page to what other LinuxKPI wlan man pages say and look like as it has been a while since I wrote it. The man page is not yet hooked up to the build on purpose as the driver is not yet enabled in the tree. Sponsored by: The FreeBSD Foundation MFC after: 3 days Reviewed by: ziaee (earlier version) Differential Revision: https://reviews.freebsd.org/D58479
There are few warnings reported by mandoc -Tlint:
bhyve_config.5:255:31: WARNING: new sentence, new line
bhyve_config.5:257:43: WARNING: new sentence, new line
bhyve_config.5:422:2: WARNING: missing section argument: Xr nm_open
bhyve_config.5:469:24: WARNING: skipping no-space macro
bhyve_config.5:483:2: WARNING: wrong number of cells: 2 columns, 4 cells
bhyve_config.5:484:2: WARNING: wrong number of cells: 2 columns, 4 cells
bhyve_config.5:541:24: WARNING: skipping no-space macro
- "new sentence, new line" is a trivial formatting fix.
- "missing section": there is actually no nm_open() manual page,
so use .Nm instead of .Xr for it.
- "no-space macro": format without .Oc and .Ns, similarly to
how it is already done in bhyve.8 for VNC addresses.
- "wrong number of cells": also a trivial fix.
Reviewed by: jhb
Sponsored by: The FreeBSD Foundation
MFC after: 3 days
Differential Revision: https://reviews.freebsd.org/D58415
`nvmecontrol power -l ...` lists the available power modes. Non-operational modes are marked with an asterisk. While here, add <device-id | namespace-id> to the "nvmecontrol power" synopsis. MFC after: 3 days Reviewed by: dab, imp, michaelo, ziaee Differential Revision: https://reviews.freebsd.org/D58480
Reviewed by: markj Sponsored by: The FreeBSD Foundation MFC after: 1 week Differential revision: https://reviews.freebsd.org/D58264
Reviewed by: emaste, mckusick Sponsored by: The FreeBSD Foundation MFC after: 3 days Differential revision: https://reviews.freebsd.org/D58592
The IPv6 socket options IPV6_JOIN_GROUP and IPV6_LEAVE_GROUP socket options are being extended to accept IPv4 multicast group addresses in the RFC 3493 IPv4-mapped address format as a convenience to application developers. Caveat this addition carefully in the newly added HISTORY section, addressing all previous review comments. Approved by: ziaee Reviewed by: ziaee, glebius PR: https://bugs.freebsd.org/bugzilla/show_bug.cgi?id=https://bugs.freebsd.org/bugzilla/show_bug.cgi?id=193246 Differential Revision: https://reviews.freebsd.org/D55382
The CTL High Availablity clustering feature allows a pair of hosts to implement transparent failover. The implementation uses a TCP connection to exchange messages. There is no authentication mechanism and the protocol itself embeds kernel pointers in the messages exchanged between HA hosts. This property (of CTL_MSG_DATAMOVE messages specifically), as well as insufficient validation of inbound messages, mean that anyone able to access a CTL HA port is able to remotely execute code on that host. Provide a warning to this effect in the CTL man page. Reported by: Ryan of Calif.io Reviewed by: ziaee, ken, mav MFC after: 3 days Sponsored by: The FreeBSD Foundation Differential Revision: https://reviews.freebsd.org/D58622
Noted and reviewed by: lwhsu Sponsored by: The FreeBSD Foundation MFC after: 1 week Differential revision: https://reviews.freebsd.org/D58666
Reviewed by: ziaee Differential Revision: https://reviews.freebsd.org/D58565
- Remove `\(em` from .Nm section as it's not valid mandoc markup. - Remove the section from the .Nm directive (it's handled under the .Dt directive). MFC after: 1 week Reported by: make manlint
Add missing commas after .Nm entries. MFC after: 1 week Reported by: make manlint
MFC after: 1 week Reported by: make manlint
MFC after: 1 week Reported by: make manlint
The shared em(4) manual page lists only the em device-node name. Document the /dev/led/igb* name as well. MFC after: 2 weeks
Reviewed by: emaste Discussed with: imp Fixes: https://cgit.freebsd.org/src/commit/?id=802c6d5d61d1 ("cdefs.h: Introduce __nonstring attribute") MFC after: 3 days Sponsored by: The FreeBSD Foundation Differential Revision: https://reviews.freebsd.org/D58804
Reviewed by: fuz Approved by: fuz (mentor) MFC after: 1 month Pull Request: https://github.com/freebsd/freebsd-src/pull/2288
MFC after: 3 days Sponsored by: fme AG
Document a few options that are currently supported but
not covered in bhyve_config(5):
- monitor
- vcpu.N.cpuset
- domains.N.{size,cpus,domain_policy}
- console (for arm64 and riscv)
MFC after: 1 week
Reviewed by: bnovkov, jhb
Sponsored by: The FreeBSD Foundation
Differential Revision: https://reviews.freebsd.org/D58399
As far as I can tell, racct(2) has never existed, not even when I added these references a decade ago. Change them as commit e9e615c88a74 did in thr_new(2). Reported by: Karlo Miličević <karlo98.m@gmail.com>
Do not duplicate the documentation already available through "sysctl -d", but tell the user where to find it. PR: https://bugs.freebsd.org/bugzilla/show_bug.cgi?id=257984 Suggested by: Felix Johnson <felix.the.red@gmail.com>
The eXpat project has changed maintainers since this section was written in 2002. Update it to reflect reality. Discussed with: Sebastian Pipping <sebastian@pipping.org> Reviewed by: bcr MFC after: 3 days Differential Revision: https://reviews.freebsd.org/D58835
Standardize driver manuals on the style used for 12 years in vt(4). This brings SYNOPSIS across all FreeBSD manual sections into harmony of meaning where where SYNOPSIS lists available options, and does not contain prose. Adjust mdoc(7) to reflect the established convention. Reviewed by: jhb Discussed with: arch@ (marc.info/?l=freebsd-arch&m=176782215606871) Differential Revision: https://reviews.freebsd.org/D54586
Point the identification LED documentation to led(4), which describes how to control /dev/led device nodes. MFC after: 2 weeks Sponsored by: BBOX.io
MFC after: 3 days
Change the path for the shosts.equiv file to consistently reflect
/etc/ssh/shosts.equiv across all manual pages.
This change stems from 35d4ccfb5576 ("Document FreeBSD defaults and
paths.")
Reviewed by: bcr, emaste
Differential Revision: https://reviews.freebsd.org/D52203
Commit 290e563166b4 changed the flag's name from dontcache to nocache. This patch fixes the man page. This is a content change. Reviewed by: kib (earlier version) MFC after: 2 weeks Differential Revision: https://reviews.freebsd.org/D58181 Fixes: https://cgit.freebsd.org/src/commit/?id=4830670a3f94 ("chflags.1: Document the new UF_DONTCACHE flag")
Fix a typo, grammar, and generally rephrase for better clarity. Fixes: https://cgit.freebsd.org/src/commit/?id=3463f02706db ("UPDATING: add an entry for [gs]etgroups") MFC after: 1 day MFC to: stable/15 Sponsored by: The FreeBSD Foundation
Reviewed by: markj Sponsored by: The FreeBSD Foundation MFC after: 1 week Differential revision: https://reviews.freebsd.org/D58586
Reviewed by: markj Sponsored by: The FreeBSD Foundation MFC after: 1 week Differential revision: https://reviews.freebsd.org/D59113
Reviewed by: markj Sponsored by: The FreeBSD Foundation MFC after: 1 week Differential revision: https://reviews.freebsd.org/D58989
MFC after: 3 days
While here, fix markup on one of the items. MFC after: 3 days
While here, tag SPDX and remove Nd overquoting. MFC after: 3 days
While here, remove Nd overquoting. MFC after: 3 days
Move the prose about support that was here to the second sentence of DESCRIPTION. Remove the note about i386 since i386 is dead in 15.0. The rest of the manual still contains details about i386 and could use some TLC by a Xen user. MFC after: 3 days (to 15 only)
Improve style for consistency with the rest of the manual and add the tunables to the search database. MFC after: 3 days Reviewed by: seuros Differential Revision: https://reviews.freebsd.org/D59165
MFC after: 3 days
While here, tag SPDX. MFC after: 3 days
MFC after: 3 days
While here, tag SPDX. MFC after: 3 days Reviewed by: kbowling Differential Revision: https://reviews.freebsd.org/D59258
While here: + tag SPDX + remove superflouous Nd quotes + add the usual missing blank comment line preceeding file MFC after: 3 days
MFC after: 3 days
MFC after: 3 days
+ tag SPDX + remove superflouous Nd quoting + fix AUTHORS email markup MFC after: 3 days
MFC after: 3 days
MFC after: 3 days
While here, remove superflouous Nd quoting. MFC after: 3 days
Reviewed by: markj Sponsored by: The FreeBSD Foundation MFC after: 1 week Differential revision: https://reviews.freebsd.org/D59132
This section will appear in the Hardware Release Note. MFC after: 3 days
MFC after: 3 days
+ Tag SPDX + Remove superflouous Nd quotes + Add "The" and "module" to introductory sentence + Remove a malformed "compact" specifier from a list + Add a missing list ending block MFC after: 3 days
MFC after: 3 days
MFC after: 3 days
MFC after: 3 days
For inclusion in the hardware release notes. MFC after: 3 days
+ Tag SPDX + Canonicalize SYNOPSIS and HARDWARE + Canonicalize SYSCTL VARIABLES, adding the sysctl to apropos + Canonicalize AUTHORS and HISTORY, which were combined + Use more traditional and consistent %d instead of less clear X MFC after: 3 days
Also tag SPDX, and switch X to %d for consistency/clarity. MFC after: 3 days
While here, remove superflouous Nd quotes and canonicalize %unit to %d. MFC after: 3 days
MFC after: 3 days
Also tag SPDX. MFC after: 3 days
Also tag SPDX. MFC after: 3 days
MFC after: 3 days
Also tag SPDX MFC after: 3 days
MFC after: 3 days
MFC after: 3 days
MFC after: 3 days
These do not belong here and can only serve to confuse search. MFC after: 3 days
MFC after: 3 days
Also tag SPDX. MFC after: 3 days
Also tag SPDX, and switch X to %d for clarity and consistency. MFC after: 3 days
+ tag SPDX + canonicalize SYNOPSIS, LOADER TUNABLES, and HARDWARE + switch X to %d for clarity and consistency MFC after: 3 days
Raising a process' priority is not permitted inside a jail, and nice(1) warns and executes the command anyway, so the service comes up at its login class priority. Not changing the date, as a commit a moment before this, one changed it already. MFC after: 1 week MFC to: stable/15
Reported by: ziaee Reviewed by: ziaee Differential Revision: https://reviews.freebsd.org/D59242
Uggg, copied this instead of using the new style. Sponsored by: Netflix
While here, document WITHOUT_DEBUG_PORTS. Approved by: dch (mentor) Approved by: kevans Approved by: ziaee Closes: https://github.com/freebsd/freebsd-src/pull/2387
- consistently use Mt request within Aq. This makes author e-mail addresses clickable in many frontends. - @freebsd.org -> @FreeBSD.org - (user@host.tld) -> Aq Mt user@host.tld Event: Berlin Hackathon 202609 MFC after: 3 days Reviewed by: ziaee Differential Revision: https://reviews.freebsd.org/D59410
+ wrap some long lines + escape some ? wildcards + no macros in width specifiers + mention the speed in HARDWARE (bumps date) + use the hyperlink macro for... the defunct LSI website... + write out a symbol heavy error message format string in mdoc MFC after: 3 days Event: Berlin Hackathon 202609
While here, s/PR#2411/NetBSD PR#2411/ in the comments for clarity. I had to go to netbsd sources to get this information. MFC after: 3 days Event: Berlin Hackathon 202609
The period here is part of the literal string in the example. Adding a space caused the example to be quoted wrongly. Instead, a trailing zero-width space keeps the linter mandoc -T happy. Reviewed by: ziaee Differential Revision: https://reviews.freebsd.org/D59353
Remove a note about "data of a different type". This was a bug that was fixed in FreeBSD 15. Instead put an exact quote from SUS that lists allowed cases of a short read with MSG_WAITALL. See discussion in D57511.
An unbraced pipe list made -n look like it could take a uid (or gid). Use braces for a single choice of name or id, and keep -u newuid / -g newgid on the name invocation only. Drop the USER/GROUP OPTIONS sentences that said -n could be a numeric id. PR: https://bugs.freebsd.org/bugzilla/show_bug.cgi?id=269193 MFC after: 3 days Reviewed by: bapt Differential Revision: https://reviews.freebsd.org/D59472
Driver-specific extension namespaces are optional. Describe their role without implying that every provider supplies one. Clarify that the VLAN count excludes membership installed implicitly by the PF without implying that an explicit VID 0 request cannot consume a reported filter. Sponsored by: BBOX.io
Reviewed by: markj Sponsored by: The FreeBSD Foundation MFC after: 1 week Differential revision: https://reviews.freebsd.org/D59349
MFC after: 3 days Event: Berlin Hackathon 202609
MFC after: 3 days Reported by: Artem Bunichev <temcbun@gmail.com> Differential Revision: https://reviews.freebsd.org/D58329
PR: https://bugs.freebsd.org/bugzilla/show_bug.cgi?id=261496 Fixes: https://cgit.freebsd.org/src/commit/?id=554442439dd0 ("Change the memory heuristics") MFC after: 3 days Reported by: wosch
Since apropos(8) uses case-insensitive regular expressions by default for both manpage names and descriptions, pfctl appears in results regardless of the parenthetical initialism. MFC after: 1 week Reported by: ziaee@
PR: https://bugs.freebsd.org/bugzilla/show_bug.cgi?id=298348 MFC after: 2 weeks
rc.conf.5: Improve NOAUTO configuration + tag SPDX - Show how to start a NOAUTOed interface - "configured" is not quite right, try to improve that MFC after: 3 days Event: EuroBSDcon Devsummit 2026 Reviewed by: adrian Discussed with: Antranig Vartanian <antranigv@freebsd.am> Differential Revision: https://reviews.freebsd.org/D59571
rc.conf.5: Fix typo Fixes: https://cgit.freebsd.org/src/commit/?id=fa7094c5b06f ("Improve NOAUTO configuration") MFC after: 3 days Reported by: Herbert J. Skuhra <herbert@gojira.at> Event: EuroBSDcon Devsummit 2026
While here, tag SPDX. Event: EuroBSDcon Devsummit 2026 MFC after: 3 days Reviewed by: wulf Differential Revision: https://reviews.freebsd.org/D59256
The resume callback also runs when suspend fails, without an intervening PCI power-state transition or configuration-space restore. Document that drivers must not assume either has occurred. MFC after: 2 weeks Sponsored by: BBOX.io
- remove old wifi drivers, these are very old and we have many more - add quick start guide to connecting to wifi Event: EuroBSDCon 2026 MFC after: 3 days Reviewed by: bz, emaste Differential Revision: https://reviews.freebsd.org/D59615
Update case synopsis, split description into paragraphs, improve markup, light wordsmithing, and import exit status line from NetBSD. While here, remove the '-' from the beginning of this document. PR: https://bugs.freebsd.org/bugzilla/show_bug.cgi?id=298329 Event: EuroBSDcon 2026 Fixes: https://cgit.freebsd.org/src/commit/?id=c9afaa63894e ("Add case statement fallthrough") Fixes: https://cgit.freebsd.org/src/commit/?id=f7a9b7fe3a8d ("Allow a left parenthesis before patterns in case blocks") Fixes: https://cgit.freebsd.org/src/commit/?id=e00e16ad7f86 ("Allow empty case/esac statements") MFC after: 3 days Reviewed by: bcr, jilles, Artem Bunichev <temcbun@gmail.com> Differential Revision: https://reviews.freebsd.org/D59518
glabel.8: clarify volume labels, metadata identifiers, and generic GEOM labels Change description of manual page to add precision. Split out filesystem, GPT and GEOM labels. Give an example for each filesystem for the program to show how to set a label. Clarify examples and use ada device names consistently. PR: https://bugs.freebsd.org/bugzilla/show_bug.cgi?id=262207 Differential Revision: https://reviews.freebsd.org/D51030
glabel.8: Fix description, improve table, tag spdx - restore missing search keyword "disk", GEOM provider is a lgoical disk - editorial nits on file system table to improve search and rendering - correct msdosfs terminology, they are FAT filesystems - use consistent spelling of filesystem Fixes: https://cgit.freebsd.org/src/commit/?id=05a2d40ac240 Event: EuroBSDcon 2026 MFC after: 3 days Discussed with: bcr, vishwin
Many of our man pages contain links to websites that are still using http. Convert them to use https. I skipped those that: - were not available anymore, - did not provide an SSL page, - had caused an SSL validation error, - were not the original link anymore (i.e. aquisition) - where the site structure had changed, causing 404s on the new site Event: EuroBSDcon 2026 Reviewed by: carlavilla (very fast, thanks for that) Differential Revision: https://reviews.freebsd.org/D59629
Event: EuroBSDcon 2026 MFC after: 3 days
Event: EuroBSDcon 2026 MFC after: 3 days
Event: EuroBSDcon 2026 MFC after: 3 days
Event: EuroBSDcon 2026 MFC after: 3 days
Sections are in all caps, and subsections are in title case. Sections also have a standard order. This manual could certainly be reformatted more to conform further to our conventions, but this is a step in the right direction. While here, remove the stray hyphen from the beginning of this document. Event: EuroBSDcon 2026 MFC after: 3 days
Event: EuroBSDcon 2026 MFC after: 3 days
Event: EuroBSDcon 2026 MFC after: 3 days
Event: EuroBSDcon 2026 MFC after: 3 days
Event: EuroBSDcon 2026 MFC after: 3 days
Event: EuroBSDcon 2026 MFC after: 3 days
Event: EuroBSDcon 2026 MFC after: 3 days
Event: EuroBSDcon 2026 MFC after: 3 days
Event: EuroBSDcon 2026 MFC after: 3 days
Event: EuroBSDcon 2026 MFC after: 3 days
Event: EuroBSDcon 2026 MFC after: 3 days
Also, add the missing blank comment line at the top of the file. Event: EuroBSDcon 2026 MFC after: 3 days
- sysutils/ataidle is no more: it was superseded by camcontrol(8). - Sort the models table alphabetically. - Use the SPDX License ID instead of the longhand licensing tort in the manpage header. - Note that the driver has been heavily modified in 15.1 and later to support additional platforms and functionality. - Trim down SYNOPSIS. MFC after: 2 weeks Differential Revision: https://reviews.freebsd.org/D59470
Netgraph document descriptions are all over the place, wordsmith them into a standard format of "%s netgraph node", trying to describe them better to enhance accessiblity of apropos results. Event: EuroBSDcon 2026 MFC after: 3 days Reviewed by: dteske, glebius Discussed with: des, dteske, glebius Differential Revision: https://reviews.freebsd.org/D59643
Document the full struct mlx5dv_context (cqe_comp_caps, sw_parsing_caps) and the mlx5dv_context_comp_mask enum in mlx5dv_query_device.3, and note that these caps are only valid when the matching comp_mask bit is asked for on input and returned on output. Sponsored by: NVidia networking MFC after: 1 month
MFC after: no Reviewed by: obiwac Differential Revision: https://reviews.freebsd.org/D59738
While here, improve HARDWARE, fix sysctl markup, and a grammatical typo. MFC after: 3 days
MFC after: 3 days
The documented VLAN and MAC filter limits are reversed. Match the defaults in the driver schema: 64 VLAN filters and 16 non-primary MAC filters per VF. MFC after: 2 weeks Sponsored by: BBOX.io Differential Revision: https://reviews.freebsd.org/D59027
Reviewed by: ziaee, kfv Differential Revision: https://reviews.freebsd.org/D59747
Reviewed by: ziaee, kfv Differential Revision: https://reviews.freebsd.org/D59746
Reviewed by: bcr MFC after: 3 days Differential Revision: https://reviews.freebsd.org/D57515
This driver is for NTT DoCoMo 3G cellular equiment, which afaict all went offline six months ago. Tidy up the entry until we can remove it. MFC after: 3 days
+ Update SYNOPSIS to the new standard format + Rename CONFIGURATION to the usual LOADER TUNABLES + Adjust tunable markup for inclusion in the search index + Editorial nit: Unwind a parenthetical MFC after: 3 days
MFC after: 3 days
MFC after: 3 days
MFC after: 3 days
+ Update SYNOPSIS to the new standard format + Give this a clearer description in apropos + Give this a clearer description in hardware note + Spinoff SYSCTL VARIABLES from HARDWARE + Pet linter for long lines and trailing punctuation + Tag SPDX MFC after: 3 days
MFC after: 3 days
MFC after: 3 days
Also add a missing Cd tag for consistency. MFC after: 3 days
MFC after: 3 days
MFC after: 3 days
+ tag SPDX + improve apropos and hardware notes context + canonicalize SYNOPSIS MFC after: 3 days
MFC after: 3 days
MFC after: 2 weeks
MFC with: f6db138f97b7 ("src.opts.mk: force the Dtrace tests ...")
I added the text when this driver was merely an import of https://github.com/Aquantia/aqtion-freebsd with patches from ports applied. nprice@ resolved issues and added support for newer cards and the note is no longer applicable. Reviewed by: adrian Sponsored by: The FreeBSD Foundation Differential Revision: https://reviews.freebsd.org/D59909
Those are FreeBSD-specific member and flag still better to be documented. Reviewed by: kib Sponsored by: Sippy Software, Inc. MFC after: 1 week Differential Revision: https://reviews.freebsd.org/D59915
Fixes: https://cgit.freebsd.org/src/commit/?id=aba599a6cc55 ("zzz: Rewrite to use new power device")
While here, fix two trivial typos. MFC after: 3 days
The standard list and status subcommands of every GEOM class emit through libxo, and so does geom -p, but only geom.8 (list, status) and gpart.8 (show) mentioned it. Mark the list and status forms with --libxo in all class manual pages, add the option description, and cross-reference xo_options(7). Also mark geom -p in geom.8. Only forms whose output already goes through libxo are marked; class-specific verbs such as gmirror dump still print directly and are left alone. Reviewed by: adrian, bcr, carlavilla, des MFC after: 3 days Differential Revision: https://reviews.freebsd.org/D59574
Sponsored by: The FreeBSD Foundation MFC after: 1 week
MFC after: 3 days Reviewed by: glebius, wulf Differential Revision: https://reviews.freebsd.org/D59869
MFC after: 3 days Reviewed by: wulf Differential Revision: https://reviews.freebsd.org/D59257
Sponsored by: The FreeBSD Foundation MFC after: 3 days
MFC after: 3 days
MFC after: 3 days
MFC after: 3 days
MFC after: 3 days
Reviewed by: ziaee, imp Differential Revision: https://reviews.freebsd.org/D59938
Reviewed by: sobomax, markj, olce Sponsored by: The FreeBSD Foundation MFC after: 1 week Differential revision: https://reviews.freebsd.org/D59910
Hardware drivers and architecture-specific code.
asmc: try PIO before MMIO to avoid false T2 detection Add hw.asmc.system-state and hw.asmc.board-id read-only sysctls to expose the T2 system state register and Mac board identifier via SMC. Try PIO access before MMIO during probe to prevent false T2 detection on Macs that happen to have something mapped at the T2 BAR address. Reviewed by: adrian Differential Revision: https://reviews.freebsd.org/D57844
asmc: add system state and board identity sysctls Add dev.asmc.0.system subtree with read-only sysctls for SMC diagnostic and identity keys: shutdown_cause (MSSD), sleep_cause (MSSP), thermal_status (MSAL), time_of_day (CLKT), power_state (MSPS), board_id (RPlt), and chip_gen (RGEN). Each sysctl is registered only if the key exists on the hardware. Reviewed by: adrian Differential Revision: https://reviews.freebsd.org/D57853
asmc: deduplicate sensor converters and cause sysctls Replace per-type spXX_to_milli() functions with a table-driven asmc_sensor_convert() that looks up the divisor by SMC type string. Reviewed by: adrian Differential Revision: https://reviews.freebsd.org/D57854
Book-E powerpc has 64-bit bus_addr_t but only a 32-bit bus_size_t. Use the right macros for maxsize and maxsegsize to fix the build. Fixes: https://cgit.freebsd.org/src/commit/?id=4bf8ce037 ("if_rge: initial import of if_rge driver from OpenBSD.") Reviewed by: adrian Differential Revision: https://reviews.freebsd.org/D57794
Depend on clknode_if.h in the module Makefile, so that it gets explicitly built for the module. Also, reduce the #if guards to only the new clock output code, and gate them on all powerpc, not just powerpc64. Fixes: https://cgit.freebsd.org/src/commit/?id=6b77d34f ("HYM8563: Add support for clock output.") Reviewed by: mmel Differential Revision: https://reviews.freebsd.org/D57795
These were added during the DPAA driver rewrite, and should not have gone in then. Remove them.
Reviewed by: mav Sponsored by: The FreeBSD Foundation MFC after: 1 week Differential revision: https://reviews.freebsd.org/D58003
PR: https://bugs.freebsd.org/bugzilla/show_bug.cgi?id=199101 MFC after: 1 week Reviewed by: imp Differential Revision: https://reviews.freebsd.org/D57929
MFC after: 1 week Reviewed by: imp Differential Revision: https://reviews.freebsd.org/D57930
Some firmware delivers the power or sleep button press that woke the system as an ordinary button press (Notify 0x80) shortly after resume, rather than as the wakeup notification (Notify 0x02) the ACPI specification requires for a button that is also a wake source. On affected machines (e.g. the Framework Laptop 12, Intel Raptor Lake-P) the power button is a control-method device behind the embedded controller. The EC latches the key press that woke the system across the sleep transition and flushes it through its normal _Qxx query path as soon as it is reinitialized on resume. The replayed press is indistinguishable from a genuine one, so the kernel honors it as a fresh suspend request and the machine suspends again immediately after waking; it cannot be kept awake with the button. The event cannot be filtered at its source: it arrives over the same EC query path that also carries legitimate events (lid, AC, thermal, battery), so suppressing the drain would lose real notifications. Instead, record the time of resume and ignore a button-initiated suspend that arrives within a short grace window of it. The timestamp is taken before DEVICE_RESUME() re-initializes the EC, so it is set before the replay can be processed on the ACPICA notify taskqueue; otherwise the replay can be evaluated before the timestamp is written and slip through. Measured from that point, the replay lands at ~600 ms across many cycles on a Framework Laptop 12, whereas a deliberate press cannot occur that quickly -- it happens well after the display is back -- so a one-second window separates the two without ignoring real presses for any perceptible time. Spec-compliant firmware reports the wake as Notify 0x02, which is handled on a different path and never reaches this check, so there is no change in behavior on such systems. The replay window is a fixed compile-time constant rather than a tunable on purpose: it tracks a hardware characteristic -- the EC's post-resume replay latency -- not a user policy, so there is no value a user would meaningfully choose. PR: https://bugs.freebsd.org/bugzilla/show_bug.cgi?id=296243 Reviewed by: adrian, imp (earlier revision), olce MFC after: 2 weeks Differential Revision: https://reviews.freebsd.org/D57712
Also, add a check in the attach method that a per-CPU structure is provided by the bus. This allows to remove such checks in multiple functions. The check cannot currently fail as all x86 CPU drivers (ACPI, legacy) provide the CPU_IVAR_PCPU instance variable, but it is safer to have it, especially as an example to other driver writers. Event: Halifax Hackathon 202606 Location: Seat 36K in AC667, still waiting for a gate at Montréal-Trudeau Sponsored by: The FreeBSD Foundation
Also, add a check in the attach method that a per-CPU structure is provided by the bus. This allows to remove such checks in multiple functions. The check cannot currently fail as all x86 CPU drivers (ACPI, legacy) provide the CPU_IVAR_PCPU instance variable, but it is safer to have it, especially as an example to other driver writers. Event: Halifax Hackathon 202606 Location: Seat 25A in AF0349, before leaving Montréal-Trudeau Sponsored by: The FreeBSD Foundation
This fixes associating to various APs. It worked fine to a FreeBSD AP (which is a wholly separate problem I'm going to need to dive into) but not to my tplink AX1800 Wifi-6 router. PR: kern/https://bugs.freebsd.org/bugzilla/show_bug.cgi?id=296503 Locally tested: * STA: Intel Centrino Advanced-N 6205 (iwn), Lenovo T420 * AP: TP-Link AX1800 wifi-6 router
M_PREPEND in the broadcast branch may call m_prepend(9) which allocates a new head mbuf and calls m_move_pkthdr(), stripping M_PKTHDR from the old mbuf. xfer->mbuf was set before M_PREPEND, so it pointed at the deheadered old mbuf. bus_dmamap_load_mbuf(9) asserts M_PKTHDR and panics. Reviewed by: zlei, adrian Differential Revision: https://reviews.freebsd.org/D57495
Use kn->kn_sdata to track the last bs->total value for each knote attached to an mmaped channel. An event is delivered only when the total byte counter has advanced by at least c->lw since the last delivery. After delivery kn_sdata is updated to the current total. Each knote tracks its own watermark independently, so multiple knotes attached to the same mmaped channel all receive events correctly. Non-mmap channels keep the existing level-triggered behavior via chn_polltrigger(). MFC after: 1 week Reviewed by: christos Differential Revision: https://reviews.freebsd.org/D57833
arm64/vmm: Add FEAT_NV2 definitions Add the definitions for the VNCR_EL2 register and all of the offsets to registers in memory relative to the page stored in VNCR_EL2. Signed-off-by: Kajetan Puchalski <kajetan.puchalski@arm.com> Reviewed by: andrew Sponsored by: Arm Ltd Differential Revision: https://reviews.freebsd.org/D56550
arm64/vmm: Use the VNCR_EL2 memory page to store guest registers Wherever possible, move the storage space for guest register values from the hypctx struct into a preallocated memory page matching the layout of the page pointed to by VNCR_EL2. This will streamline implementing support for nested virtualization, but the implementation itself is not reliant on the presence of nested virtualization architecture features. Signed-off-by: Kajetan Puchalski <kajetan.puchalski@arm.com> Reviewed by: andrew Sponsored by: Arm Ltd Differential Revision: https://reviews.freebsd.org/D56551
arm64/vmm: Store non-VNCR registers in an array Move non-VNCR EL0 and EL1 registers into a dedicated array inside of hypctx. This enables uniform accesses to both VNCR and non-VNCR guest register state through hypctx_[read|write]_sys_reg(). The accessors are _not_ used for non-VNCR EL2 registers in order to create a clear separation between guest-visible and guest-invisible register state. Signed-off-by: Kajetan Puchalski <kajetan.puchalski@arm.com> Reviewed by: andrew Sponsored by: Arm Ltd Differential Revision: https://reviews.freebsd.org/D56552
arm64/vmm: Refactor vmm_hyp.c Refactor vmm_hyp.c to split register reload logic by type of register, streamline the implementation and improve readability. Signed-off-by: Kajetan Puchalski <kajetan.puchalski@arm.com> Reviewed by: andrew Sponsored by: Arm Ltd Differential Revision: https://reviews.freebsd.org/D56553
arm64/vmm: Move vttbr_el2 & vtimer into struct hypctx Move vttbr_el2 & vtimer from struct hyp into struct hypctx to streamline the logic and handle them in the same way as other *_el2 registers are already being handled. Signed-off-by: Kajetan Puchalski <kajetan.puchalski@arm.com> Reviewed by: andrew Sponsored by: Arm Ltd Differential Revision: https://reviews.freebsd.org/D56554
arm64/vmm: Move host-side EL2 regs into sys_regs Move EL2 host registers that are not visible to the guest into hypctx->sys_regs. Prefix them with HOST_ to distinguish from EL2 registers which are part of the guest's own state (e.g. in VNCR). Signed-off-by: Kajetan Puchalski <kajetan.puchalski@arm.com> Reviewed by: andrew Sponsored by: Arm Ltd Differential Revision: https://reviews.freebsd.org/D56555
arm64/vmm: Make remaining registers use hypctx_*_sys_reg Move vgic, timer and trapframe registers into sys_regs to handle them in the same way as all the other registers. Signed-off-by: Kajetan Puchalski <kajetan.puchalski@arm.com> Reviewed by: andrew Sponsored by: Arm Ltd Differential Revision: https://reviews.freebsd.org/D56556
Fix tid_set_busy() for when `pmap` is NULL. Obviously a NULL pointer cannot be correctly used, so I'm not sure how it worked in testing on 64-bit.
When the TID rolls over on a given CPU, simply flash-invalidate the TLB instead of walking the TLB to only invalidate the repurposed TID. Walking 256 entries is expensive, and we'll likely be inserting a bunch new ones anyway in the new environment, since 256 really only handles 1MB of storage, so the likelihood of other mappings continuing to exist in the TLB when their thread owner is scheduled again is very very small.
DEVX event notifier returned true for the command-completion and page-request events. This is causing mlx5_eq_int() to skip the core EQ handler, so the firmware command interface and the page supply stop being serviced and the device wedges. This commit also make notifier registration and dispatch safe against the EQ interrupt running concurrently: publish the table pointer before the callback and load it with acquire semantics. run the callback under RCU, and drain it with synchronize_rcu() on teardown. Otherwise the interrupt handler could observe a half-initialized notifier or race with cleanup. Reviewed by: kib Tested by: Wafa Hamzah <wafah@nvidia.com> Sponsored by: Nvidia networking MFC after: 1 month
Import Linux upstream commits a8b92ca1b0e5ce620e425e9d2f89ce44f1a82a82 and c59450c463695a016e823175bac421cff219935d. The DEVX object and method definitions were already present, but nothing pointed ib_device.driver_def at them. ibcore therefore never merged them into the uverbs uapi tree and every DEVX ioctl came back as EPROTONOSUPPORT. Reviewed by: kib Tested by: Wafa Hamzah <wafah@nvidia.com> Sponsored by: Nvidia networking MFC after: 1 month
Import Linux upstream commit 342ee59de98a2ecdf15a46849a2534e7c808eb1f. The dynamic UAR object was declared in the ABI headers but had no handler, so the ioctl was rejected and dynamic-UAR contexts could not allocate a doorbell UAR at all. Implement the alloc and destroy methods following the upstream driver: grab a UAR stamped with the caller's DEVX uid, expose it to user space through an rdma_user_mmap entry (write-combining or non-cached as requested), and free it on destroy. Reviewed by: kib Tested by: Wafa Hamzah <wafah@nvidia.com> Sponsored by: Nvidia networking MFC after: 1 month
A firmware object owned by a DEVX uid may only reference resources owned
by the same uid or ones explicitly marked as shared. Completion EQs
were created with uid 0, so a CQ owned by a DEVX uid could not attach to
its EQ and CREATE_CQ failed with "bad resource".
Create completion EQs with MLX5_SHARED_RESOURCE_UID on devices that
support user contexts, so uid-owned CQs can use them.
The code follows the Linux commit d2c8a1554c10d5e0443b1f97f480d7dacd55cf55
("IB/mlx5: Enable UAR to have DevX UID").
Reviewed by: kib
Tested by: Wafa Hamzah <wafah@nvidia.com>
Sponsored by: Nvidia networking
MFC after: 1 month
mlx5ib: allocate IB queue counters as a shared resource
A QP owned by a DEVX uid references the port's queue counter. The
counter was allocated with uid 0, so RST2INIT_QP on a uid-owned QP
failed with "bad resource state".
Allocate and free the IB queue counters directly and, on devices that
support user contexts, stamp them with MLX5_SHARED_RESOURCE_UID so
uid-owned QPs can use them.
The code follows the Linux commit d2c8a1554c10d5e0443b1f97f480d7dacd55cf55
("IB/mlx5: Enable UAR to have DevX UID").
Reviewed by: kib
Tested by: Wafa Hamzah <wafah@nvidia.com>
Sponsored by: Nvidia networking
MFC after: 1 month
mlx5ib: encode dynamic UAR mmap offsets in the reserved command range The UAR ioctl handed user space a raw mmap offset, so the first dynamic UAR landed at page offset 0. mlx5_ib_mmap() decodes offset 0 as the legacy regular-page command and routed the mapping through the old bfreg path, which rejects dynamic-UAR contexts, so mmap() failed with EINVAL and mlx5dv_devx_alloc_uar() returned NULL. Follow the upstream scheme: reserve the mmap command range [9, 255] for rdma_user_mmap entries and return command-encoded offsets, so the dynamic-UAR mappings decode to the intended mlx5_ib_mmap() path. Reviewed by: kib Tested by: Wafa Hamzah <wafah@nvidia.com> Sponsored by: Nvidia networking MFC after: 1 month
mlx5ib: advertise write-combining support for dynamic BlueFlame UARs Import Linux upstream commit 1f3db161881b7e21efb149e0ae8152b79a571a8f. dev->wc_support was never set, so it was always false and the UAR ioctl refused BlueFlame (write-combining) UAR allocations with EOPNOTSUPP. That breaks QP creation in pure dynamic-UAR mode, where user space asks for a BF doorbell UAR. Reviewed by: kib Tested by: Wafa Hamzah <wafah@nvidia.com> Sponsored by: Nvidia networking MFC after: 1 month
mlx5: pass the full EQE to the DEVX event notifier The DEVX event notifier and its helpers expect a full struct mlx5_eqe and read eqe->data from it, but mlx5_eq_int() passed &eqe->data, so the data offset was applied twice. Reviewed by: kib Tested by: Wafa Hamzah <wafah@nvidia.com> Sponsored by: Nvidia networking MFC after: 1 month
mlx5: guard against a NULL CQ event handler in mlx5_cq_event() DEVX and mlx5en created CQs are registered without an asynchronous event handler (mcq.event is NULL). An asynchronous CQ_ERROR event for such a CQ made mlx5_cq_event() call through a NULL pointer and panic. Reviewed by: kib Tested by: Wafa Hamzah <wafah@nvidia.com> Sponsored by: Nvidia networking MFC after: 1 month
mlx5: propagate the DEVX uid through SRQ create and destroy The SRQ command builders never stamped the owning DEVX uid into the firmware CREATE_SRQ/CREATE_RMP/CREATE_XRC_SRQ commands, so a basic SRQ was always created with uid 0. Every modern libmlx5 context runs with a DEVX uid, and the QPs that reference the SRQ carry that uid, so firmware rejected CREATE_QP with "bad resource": a uid-owned QP may not reference a uid-0 SRQ. Reviewed by: kib Tested by: Wafa Hamzah <wafah@nvidia.com> Sponsored by: Nvidia networking MFC after: 1 month
The thermal interrupt is initially masked. Thermal interrupt handling is enabled by calling lapic_enable_thermal(), which installs a (single) handler. [olce: Wrote the commit message.] Reviewed by: kib, olce MFC after: 2 weeks Differential Revision: https://reviews.freebsd.org/D44454
acpi: Add a pseudo-bus for APEI devices to manage resources Different APEI tables can reuse the same registers (and sometimes different views of the same register, e.g. 32- vs 64-bit mappings of the same register). To enable this sharing, apei0 now acts as a bus device managing a pool of allocated resources and handing out mappings to child devices which handle individual tables. Most of the previous apei(4) driver has been moved into a new hest0 device that is a child of apei0. Reviewed by: gallatin Sponsored by: Netflix Differential Revision: https://reviews.freebsd.org/D58024
acpi: fix instant panic in hest_attach() Since now there is a pseudo-bus between our device and acpi0, we need to go deeper. Fixes: https://cgit.freebsd.org/src/commit/?id=9313f6b01485ad9a0b7cc59b459f5714533587c3
This driver parses the ACPI EINJ table and builds a list of instructions associated with known actions. It then exports ioctls to fetch the set of supported errors and inject system errors by executing specific sequences of actions. This can be used to test error reporting facilities for events such as ECC errors. Reviewed by: gallatin Sponsored by: Netflix Differential Revision: https://reviews.freebsd.org/D58025
Fixed the post-LPS delay from 500us to the IEEE 1394a-2000 s6.1 mandated 10ms ceiling. Handled PHY_INT by clearing W1C status bits in register 5 (masked ISBR to avoid spurious bus resets). Added a SID timeout callout that recovers the state machine when a remote device fails to complete self-ID. Fixed FW_PHY_SPD operator precedence and gated noisy messages behind bootverbose/firewire_debug. Reviewed by: adrian Differential Revision: https://reviews.freebsd.org/D58033
fwcam: add dynamic resolution and frame rate support Read V_MODE_INQ and V_RATE_INQ registers for all supported formats during probe, caching the camera's actual capabilities. Use these to validate SMODE ioctl requests before writing to the camera. Writing an unsupported combination caused the camera to stop responding, requiring a physical power cycle. Tested with: Apple iSight (external FireWire) Reviewed by: adrian Differential Revision: https://reviews.freebsd.org/D58090
fwcam: set ISO speed from device link speed iso_speed was never initialized, defaulting to S100 regardless of the camera's actual link speed. Some cameras firmwares reject ISO_EN when the speed field in the ISO_CHANNEL register does not match their capabilities. Reviewed by: adrian Differential Revision: https://reviews.freebsd.org/D58091
fwcam: write video mode registers before enabling ISO streaming The IIDC spec (s3.1) requires the video mode to be programmed before ISO enable. Without this, cameras that power up with invalid default mode/rate combinations reject the ISO_EN write. This can happen when the firmware of teh camera is outdated or vendor never updated it. Differential Revision: https://reviews.freebsd.org/D58092
fwcamctl provides userland access to /dev/fwcam0. Supported subcommands: info (camera state, format, mode, rate, features), snap (capture a frame as PPM), mode (set format/mode/rate), and feat (get/set camera feature registers). snap converts YUV422, YUV411, YUV444, RGB8, and Mono8 pixel formats to RGB24 PPM with no external dependencies. A configurable frame skip (default 5) allows auto-exposure and auto-white-balance to settle before capture. (from adrian - yes, I've successfully captured images from an Apple isight camera on firewire with this tool and in-tree support.) Reviewed by: adrian Differential Revision: https://reviews.freebsd.org/D57914
Sponsored by; The FreeBSD Foundation MFC after: 1 week
Sponsored by: The FreeBSD Foundation MFC after: 1 week
Moved ISO start to first usage. Opening the device now only validates state and increments the open count, allowing info queries and mode changes without starting the camera. ISO streaming begins on demand when userland first reads frame data. This avoid the camera led to turn-on at attach. Reviewed by: adrian Differential Revision: https://reviews.freebsd.org/D58100
Some IIDC cameras power down the sensor when inactive (e.g. lens cover closed) and reject ISO enable with EIO. Re-power the camera and retry once before failing. Reviewed by: adrian Differential Revision: https://reviews.freebsd.org/D58101
Expose audio capture from Apple FireWire devices as a standard pcm(4)/dsp(4) device via the newpcm framework. (adrian: I've tested this on an isight camera and looped it back to USB speakers via "sox -t oss /dev/dsp3 -t oss /dev/dsp4") Reviewed by: adrian Differential Revision: https://reviews.freebsd.org/D58109
This driver only reports the RFKILL button presses. This is needed for the "airplane" key on some Framework laptops. Reviewed by: wulf, ziaee Event: Halifax Hackathon 202606 Location: vishwin@'s car Co-authored-by: Daniel Shaefer Sponsored by: Framework Computer Inc Sponsored by: The FreeBSD Foundation Differential Revision: https://reviews.freebsd.org/D57838
MFC after: 1 week Sponsored by: The FreeBSD Foundation
USB vendor:product 184f:0051 Sponsored by: The FreeBSD Foundation Differential Revision: https://reviews.freebsd.org/D56794
This can't be a loadable module, so add it to MINIMAL Sponsored by: Netflix Differential Revision: https://reviews.freebsd.org/D58067
This makes the code slightly more compact and easier to read. No functional change intended. Reviewed by: bnovkov MFC after: 2 weeks Sponsored by: The FreeBSD Foundation Differential Revision: https://reviews.freebsd.org/D58110
The thermal LVT slot does not necessarily exist. According to Intel's Software Developers Manual, for Intel processors supporting 64-bit operation (amd64), probably even the earliest ones should have a local APIC with such a slot (the slot was introduced with Pentium 4 and Xeon processors according to the manual, and the 64-bit implementation in some later versions of them). AMD's Architecture Programmer's Manual also seems to imply that all AMD processors supporting amd64 should have the slot too. So this change may not be needed when i386's code is dropped, but it does not hurt to have it, and it might ease possible MFCs. Change the signature of lapic_enable_thermal() so that it can report failure (if there is no local APIC or if there is no thermal LVT slot). Reviewed by: bnovkov, kib MFC after: 2 weeks Sponsored by: The FreeBSD Foundation Differential Revision: https://reviews.freebsd.org/D58086
Differential Revision: https://reviews.freebsd.org/D56923 Reviewed by: mhorne
This change implements the equivalent of the amd64-specific 'show pte' ddb command used to dump the page table entries associated with a specific virtual address. Differential Revision: https://reviews.freebsd.org/D56924 Reviewed by: mhorne
Spurious page faults caused by cached invalid entries may occur when starting APs and potentially panic the kernel if we're running in a non-sleepable context. Fix this avoidable panic by flushing the TLB after the AP is released. Differential Revision: https://reviews.freebsd.org/D57003 Reviewed by: markj
The Privileged ISA specification permits caching of invalid PTEs 12.2.1. Supervisor Memory-Management Fence Instruction), which may result in a spurious page fault. Such faults are handled by 'pmap_fault' which locks the kernel pmap before inspecting and possibly updating the offending L2 entry. Unfortunately, spurious faults may also occur when we're already holding the kernel_pmap lock or running in a critical section, where any attempt to grab the pmap lock will result in a kernel panic. Fix this avoidable panic by performing a lockless lookup to determine whether a valid kernel mapping exits and flushing appropriate TLB entry. Differential Revision: https://reviews.freebsd.org/D56925 Reviewed by: jrtc27, mhorne, markj
Using cpu_get_pcpuid() directly or having a CPU ID cache does not really make any significant difference. With cache: Less function calls, less space on stack, but an additional allocation in the softc, who stays permanently. Without cache: Some function calls, but one less slot in the softc, and no data duplication (but that info never changes). The main reason for this change is to reduce conflicts with some work-in-progress by aokblast@. While here, move the check that a per-CPU structure is provided by the bus from the attach to the probe method, as it is already used by hwpstate_probe_pstate() there. Reviewed by: aokblast Sponsored by: The FreeBSD Foundation
To minimize the diff with hwpstate_intel(4). See previous commit there for the rationale. Reviewed by: aokblast Sponsored by: The FreeBSD Foundation
KVM does not always use 0x40000000 as its CPUID base. For example, QEMU adds a 0x100 offset when nested virtualization is detected and the host exposes Hyper-V enlightenment hints. To accommodate this behavior, switch the detection logic to use the CPUID leaf returned by do_cpuid(), making the implementation more flexible. See: https://github.com/qemu/qemu/blob/master/target/i386/kvm/kvm.c#L2300 Reviewed by: kib MFC after: 2 weeks Sponsored by: The FreeBSD Foundation Differential Revision: https://reviews.freebsd.org/D58146
Presumably surfaced by -fstack-protector-strong, rk8xx_settime was triggering SSP when ntpd set the time on the RockPro64, at the very least. A minor oops meant that the weeks mask was getting tossed into the wrong field, and the mask was never populated. The mask is 0x7 for all three of these, thus overflowing the `data` array in settime by one byte. PR: https://bugs.freebsd.org/bugzilla/show_bug.cgi?id=296719 Reported by: jsm, "Tenkawa" on Discord Reviewed by: mmel Differential Revision: https://reviews.freebsd.org/D58182
Add fwdv(4) driver for DV video capture from FireWire camcorders using AV/C protocol and isochronous streaming. Supports AV/C tape transport commands (play, stop, ff, rewind, pause, record, eject) with NTSC/PAL auto-detection and read(2) interface for frame capture. Reviewed by: adrian Differential Revision: https://reviews.freebsd.org/D58122
Using if_getflags() to check IFF_DRV_RUNNING is wrong; if_getdrvflags() is required. This issue resulted in the multicast filter not being updated. This was an oversight by me in my initial port. Thanks to danilo@ for reporting it and Oleg <oleglelchuk@gmail.com> for the fix. PR: kern/https://bugs.freebsd.org/bugzilla/show_bug.cgi?id=295176
The Wacom ExpressKey Remote (ACK-411050) is a wireless button pad
with 18 programmable buttons and a touch ring, used as a companion
device with Wacom tablets.
It communicates via a USB wireless receiver (0x056a:0x0331) using a
vendor-specific HID report (ID 0x11).
This driver exposes the device via evdev:
- 18 buttons: BTN_0–BTN_Z, BTN_BASE, BTN_BASE2
- Touch ring position via ABS_WHEEL (0–71; reports 0 on release)
- Pad activity marker via ABS_MISC (set to 15/PAD_DEVICE_ID when
any input is active, 0 when idle that matches Linux wacom driver
convention)
- Remote serial number via MSC_SERIAL (for userland per-remote
identification)
Battery level, charging state, and touch ring mode (3 LEDs, values 0–2)
are exposed as per-device sysctls (dev.hidwacom.0.battery, .charging,
.ring_mode) rather than overloading evdev misc codes. The ring mode
sysctl is preserved across device idle periods.
Protocol was decoded from USB traffic analysis and cross-referenced
against the Linux wacom_remote_irq() implementation in
drivers/hid/wacom_wac.c.
Reviewed by: adrian
Differential Revision: https://reviews.freebsd.org/D56729
Discussed with: ziaee
Currently, devd emits events for external adapters only. Send Netgraph init/disconnect events to devd so the internal adapter's state could be asserted from userland. (adrian - indentation changes.) Signed-off-by: Kirill Orlov (-k) <slowdive@me.com> Reviewed-by: adrian, imp Pull-Request: https://github.com/freebsd/freebsd-src/pull/2196
Matches tcpdump naming, but without getting more intense as you add more -t. This slightly reduces the post-processing needed on usbdump output to diff two transactions. Reviewed by: adrian Differential Revision: https://reviews.freebsd.org/D58196
Introduce fdt_ether_get_addr() in fdt_common.c/h that tries standard DT properties in the correct order and falls back to a random address when needed. This should be used by ethernet drivers instead of open-coding the same logic. MFC after: 2 weeks Reviewed by: mhorne, adrian, bz, jrtc27 Differential Revision: https://reviews.freebsd.org/D58104
When compiled without 'options RSS', the ena driver created taskqueues using taskqueue_start_threads_cpuset passing a mask value of NULL, both in the ena_setup_tx_resources path (for enqueues) and in the ena_create_io_queues path (for the completion-processing). In the default configuration, on most EC2 instances, this results in taskqueues running in the right NUMA domain, but only by accident; in non-default configurations (e.g. with with multiple EBS volumes attached and associated NVMe taskqueues) the taskqueues may land in the wrong NUMA domain even on instance types where the one-EBS-one-ENA case produces the desired results. Set (struct ena_que)->domain and use that to inform the choice of CPU sets. On a c8gn.48xlarge EC2 instance this doubles throughput on a 32-TCP-stream benchmark. Reviewed by: akiyano MFC after: 7 days Sponsored by: Amazon Differential Revision: https://reviews.freebsd.org/D57918
In the DEVX_SUBSCRIBE_EVENT handler the eventfd path can fail and "goto err" before the subscription's xa keys and ev_file have been set; they are still zeroed from kzalloc(). The cleanup then looks up a level-1 xa entry with key 0, gets NULL, and faults dereferencing it. Initialize the fields the cleanup path relies on right after the subscription is allocated, before it is linked and before the fallible fdget(), so a later failure unwinds cleanly. Reviewed by: kib Sponsored by: Nvidia networking MFC after: 1 month
The DEVX_SUBSCRIBE_EVENT redirect path resolved the user's eventfd with fdget(), which on FreeBSD only finds LinuxKPI files. rdma-core creates the eventfd with the native FreeBSD eventfd(2), so the lookup failed and subscription returned EBADF; the delivery side likewise assumed a LinuxKPI-pollable file. Use the LinuxKPI eventfd_ctx API instead: eventfd_ctx_fdget() resolves the native eventfd, eventfd_signal() notifies it, and eventfd_ctx_put() releases it. DEVX async events can then be delivered through a redirect eventfd. Reviewed by: kib Sponsored by: Nvidia networking MFC after: 1 month
Fixes: https://cgit.freebsd.org/src/commit/?id=fc9dc8482396 ("snd_uaudio: Lock usbd_transfer_start() in uaudio_mixer_ctl_set()") PR: https://bugs.freebsd.org/bugzilla/show_bug.cgi?id=296682 Sponsored by: The FreeBSD Foundation MFC after: 3 days
virtio: Add feature bit definitions up to VirtIO v1.3 Signed-off-by: Faraz Vahedi <kfv@kfv.io> Reviewed-by: ngie Pull-Request: https://github.com/freebsd/freebsd-src/pull/2319
virtio: Report feature masks on negotiation failure Signed-off-by: Faraz Vahedi <kfv@kfv.io> Reviewed-by: ngie Pull-Request: https://github.com/freebsd/freebsd-src/pull/2319
virtio: Accept VIRTIO_F_RING_RESET in the modern PCI transport Accept per-virtqueue reset when the device offers it, alongside the V1 flag. Negotiating the feature merely permits the use of per-virtqueue reset and imposes no obligation on a driver that never uses it, while refusing capability-only transport features can make strict devices reject the feature set altogether. No functional change on hosts that do not offer RING_RESET. Signed-off-by: Faraz Vahedi <kfv@kfv.io> Reviewed-by: ngie Pull-Request: https://github.com/freebsd/freebsd-src/pull/2319
This is to prevent child drivers from using the features returned by previous drivers (in an arbitrary order). None of the existing ones do that, so this is purely defensive. MFC after: 2 weeks Sponsored by: The FreeBSD Foundation
This is needed for VM_PHYS_TO_PAGE() to work, which is needed for pmap_map_io_transient() to work, which is needed for uiomove_fromphys() to work. PR: https://bugs.freebsd.org/bugzilla/show_bug.cgi?id=296348 Reported and tested by: Anton Saietskii <vsasjason@gmail.com> Reviewed by: markj Sponsored by: The FreeBSD Foundation MFC after: 1 week Differential revision: https://reviews.freebsd.org/D58274
As RX processing is heavier than TX completions processing, swap the order and process TX completions first, in order to avoid starving the completions and causing potential missing TX completions. Submitted by: Ofir Tabachnik <ofirt@amazon.com> MFC after: 2 weeks Sponsored by: Amazon, Inc. Reviewed by: cperciva Differential Revision: https://reviews.freebsd.org/D58239
Move per-packet counter_enter/counter_exit pairs out of the RX processing loop and batch them into a single update after the loop completes. Previously, each received packet triggered two separate counter_enter/counter_exit blocks -- one for bytes and one for packet count. This commit accumulates totals in local variables and updates all four counters (ring and hw stats for both packets and bytes) in a single counter_enter/counter_exit block after the loop. Also move the stats update to after the refill and LRO flush so that the error path (goto update_stats) and the normal path converge at the same label, avoiding code duplication. Submitted by: David Arinzon <darinzon@amazon.com> MFC after: 2 weeks Sponsored by: Amazon, Inc. Reviewed by: cperciva Differential Revision: https://reviews.freebsd.org/D58240
Sporadic 'Found a Tx that wasn't completed on time' warnings appear
under sustained TX load, always reporting '1 msecs since last cleanup'
despite the 5-second timeout threshold.
The per-packet TX timestamp uses struct bintime (128 bits: two 64-bit
fields sec and frac) which is read and written non-atomically. A race
exists between the missing TX completion check
(check_missing_comp_in_tx_queue reading the timestamp) and the TX
submit path or cleanup path writing it on another CPU. Since the two
fields are not updated atomically, the check can observe a partially
written timestamp - one field from the old value and one from the new.
This can produce a timestamp with {sec=0, frac=valid}, causing the
check to compute a time offset equal to system uptime and falsely
exceeding the 5-second timeout.
Confirmed by instrumentation showing all occurrences had sec=0 with
valid frac/mbuf, cleanup_running=0, and ticks==last_cleanup_ticks.
Replace struct bintime with sbintime_t (a single 64-bit value) for
tx_buf->timestamp. An aligned 64-bit store/load cannot be torn on
64-bit architectures. Additionally, snapshot the timestamp into a
local variable in the check path to prevent a read-then-read race
where the timestamp could be zeroed between the zero-check and the
offset calculation.
Testing:
On m6i.large (FreeBSD 15.0-RELEASE-p6 amd64, 2 IO queues), two
instances with MTU 1500. Ran iperf -P 20 -u -b 320kpps (CPU
saturated at ~7 Gbps aggregate).
Without the fix: 8 warnings in 6 hours (first at ~72 min).
With the fix: 0 warnings after 20+ hours under identical conditions.
Fixes: https://cgit.freebsd.org/src/commit/?id=9b8d05b8ac78 ("Add support for Amazon Elastic Network Adapter (ENA) NIC")
Submitted by: Gilad Ben Yakov <giladben@amazon.com>
MFC after: 2 weeks
Sponsored by: Amazon, Inc.
Reviewed by: cperciva
Differential Revision: https://reviews.freebsd.org/D58241
Bug Fixes: * Fix false 'missing TX completions' warnings due to timestamp race * Put taskqueues into correct NUMA domain if !RSS Minor Changes: * Batch RX statistics updates * Swap RX/TX completions cleanup order Submitted by: Arthur Kiyanovski <akiyano@amazon.com> MFC after: 2 weeks Sponsored by: Amazon, Inc. Reviewed by: cperciva Differential Revision: https://reviews.freebsd.org/D58242
Added structure to allow multiple device to attach to the same driver. Also removed the deprecation warning from the man page. Differential Revision: https://reviews.freebsd.org/D58201 Reviewed by: adrian
Migrated fwcam to use per-unit-directory child device Differential Revision: https://reviews.freebsd.org/D58202 Reviewed by: adrian
Migrated fwisound to use per-unit-directory child device Differential Revision: https://reviews.freebsd.org/D58203 Reviewed by: adrian
Migrated fwdv to use per-unit-directory child device Reviewed by: adrian Differential Revision: https://reviews.freebsd.org/D58204
SPL is a no-op on amd64. Real locking is already handled by fc_mtx and per-driver mutexes. Reviewed by: imp Differential Revision: https://reviews.freebsd.org/D58210
Otherwise we try to disable the wrong IRQ. Fixes: https://cgit.freebsd.org/src/commit/?id=47e073941f4e ("Import the kernel parts of bhyve/arm64") MFC after: 1 week
Now that IRQs can properly be disabled by GICD_ICENABLERn, an EOI for a disabled IRQ ends up being lost, since we don't assign it to a list register and don't enable maintenance interrupts for such cases. As a result, we keep the IRQ active, which stops it from ever being delivered again (which would be true even if we supported the active and pending state). Keep disabled but active IRQs around in list registers so we can see the EOI having taken place in a future sync (noting that since we already don't create list registers in active and pending state there are no concerns with causing a disabled IRQ to be delivered). Fixes: https://cgit.freebsd.org/src/commit/?id=47e073941f4e ("Import the kernel parts of bhyve/arm64") MFC after: 1 week
dpaa2_ni_init() only enabled the DPNI object; it never pushed the
promiscuous/allmulti state or the multicast filter table to the MC
firmware. The SIOCSIFFLAGS handler ignores flag changes that arrive
while the interface is down, yet still latches them into sc->if_flags,
so a promiscuous mode request made before the first up was silently
lost and could never be applied afterwards: the up path runs
dpaa2_ni_init(), which did not read the flags, and every later
SIOCSIFFLAGS compares against the already-latched value and sees no
change.
This is exactly what happens when if_bridge adds a dpni member while
the dpni is still down, e.g. rc.conf's
create_args_bridge0="... addm dpni0"
running at bridge clone time, before ifconfig_dpni0="up" is processed.
bridge_ioctl_add() puts the member into promiscuous mode at addm time;
the request never reaches the firmware, so the DPNI continues to
hardware-filter unicast destined to other MACs. ifconfig still
reports PROMISC (a stack-level flag), which makes the failure
invisible: the host stays reachable only via the DPNI's own MAC
address (e.g. with net.link.bridge.inherit_mac=1), while bridged
epair/vnet jail traffic is silently dropped on RX.
Reapply both pieces of administrative state after enabling the DPNI,
as other NIC drivers do in their init path. This also restores
multicast memberships joined while the interface was down.
PR: https://bugs.freebsd.org/bugzilla/show_bug.cgi?id=292006
Reported by: jhibbits
Signed-off-by: Nick Price <nick@spun.io>
Reviewed by: jhibbits
Differential Revision: https://reviews.freebsd.org/D58330
This fixes a build break for i386. Reviewed by: kib, olce, Koine Yuusuke <koinec@yahoo.co.jp> Fixes: https://cgit.freebsd.org/src/commit/?id=87ba088fa310 ("x86/local_apic.c: Add support for installing a thermal interrupt handler") Differential Revision: https://reviews.freebsd.org/D58332
hwpstate_intel: Fix i386 build Reviewed by: olce Fixes: https://cgit.freebsd.org/src/commit/?id=7b26353a59d6 MFC after: 3 days Sponsored by: The FreeBSD Foundation Differential Revision: https://reviews.freebsd.org/D58208
hwpstate_intel: Minimize ifdef for i386 build Reported by: jrtc27 Fixes: https://cgit.freebsd.org/src/commit/?id=bdc0f7678257 MFC after: 3 days Sponsored by: The FreeBSD Foundation
aq(4): expand and correct offloads, fix VLAN/multicast filtering Advertise the offloads the hardware already performs, correct the TX descriptor's L3 family selection, and correct the VLAN and multicast receive-filter paths. Offloads: advertise IFCAP_HWCSUM_IPV6 (adding CSUM_IP6_TCP/UDP/TSO to isc_tx_csum_flags) and IFCAP_VLAN_HWTSO, and enable the RX outer (S-VLAN) tag parse mode in aq_hw_offload_set(). TX descriptor L3 family: aq_setup_offloads() derived tx_desc_cmd_ipv4 from CSUM_IP|CSUM_TSO, but CSUM_TSO is (CSUM_IP_TSO|CSUM_IP6_TSO) and tcp_output() sets both bits without regard to address family, so an IPv6 TSO frame matched on CSUM_IP_TSO and went out with the IPv4 header-checksum command set on a frame that carries no IPv4 header. The checksum flags cannot distinguish the family; key the bit off IPI_TX_IPV4 instead, which iflib derives from the parsed ethertype, as the IPI_TX_INTR test below it already does. Plain IPv6 checksum offload was unaffected, as CSUM_IP6_TCP alone never matched the mask. RX VLAN tag stripping: ring init hardwired hardware tag stripping off while the RX path still set M_VLANTAG and the writeback tag for every tagged frame, so a tagged frame arrived with the tag in line while the mbuf claimed it stripped and ether_demux() parsed four bytes short of the payload. Program per-ring stripping from IFCAP_VLAN_HWTAGGING and set M_VLANTAG only under the same capability, so the two states stay coherent. VLAN filter and promiscuous edge cases: filter only when 1..16 VLANs are registered -- with none (or more than the 16 the table holds) fall back to VLAN-promiscuous and pass all tags, rather than dropping every tagged frame against an empty filter table; and keep VLAN-promiscuous set whenever the interface is IFF_PROMISC, so adding or removing a VLAN under promisc does not clear it and start dropping tagged frames. Multicast reconcile: ifdi_multi_set is declarative, but aq_if_multi_set() only added -- shrinking the list left accept-all-multicast latched or stale exact slots enabled, defeating hardware multicast filtering until a reinit. Clear the exact slots before reprogramming the current list, and always drive accept-all-multicast from the current state so a shrink clears it. Reviewed by: adrian Differential Revision: https://reviews.freebsd.org/D58145
aq(4): drop errored RX frames instead of resetting the interface aq_isc_rxd_pkt_get() returned EBADMSG when a receive descriptor's MAC/receive-error bit (rx_stat bit 0) was set. iflib treats any error from isc_rxd_pkt_get() as a fatal ring fault and answers with IFC_DO_RESET -- a full interface reinitialization. A per-frame receive error is not a ring fault: on a marginal link or cable the Atlantic delivers errored frames continuously, so each one triggered another reset and the interface reset-stormed itself into carrying no traffic instead of merely dropping the bad frames. The Atlantic delivers errored frames to the host by design (Linux drops them in software via buff->is_error), and iflib offers no per-frame error return that isn't a reset. Follow the vmxnet3 model: on a receive error zero the fragment lengths and return success. iflib then discards the packet (assemble_segments() excludes zero-length fragments) while still recycling the descriptors through the refill path -- no reset. Also drop frames flagged with an RX-DMA fault (rdm_err), not just the MAC-error bit; and keep iri_len non-zero on that drop path, since iflib asserts iri_len != 0. The genuinely structural errors -- more segments than isc_rx_nsegments, or a pkt_len inconsistent with the descriptor count -- still return EBADMSG, since those indicate a confused ring where a reset is the right recovery. Reviewed by: adrian Differential Revision: https://reviews.freebsd.org/D58136
aq(4): honor the kernel RSS policy and add a TX traffic-class helper Align RX steering with the kernel RSS framework and factor out the active-traffic-class count. RSS key and indirection table: on an options RSS kernel the stack owns a canonical hash key and a hash-to-bucket indirection table binding each bucket to a CPU. aq programmed a random arc4rand() key and a plain i % rss_qs table, so the hash it stamped in iri_flowid and the queue it steered a flow to did not match the CPU the stack chose -- defeating RSS affinity. Under #ifdef RSS take the key from rss_getkey() and each entry from rss_get_indirection_to_bucket(), as e1000/ixgbe/ixl do; the non-RSS build keeps the random key and round-robin table. RSS hash-type policy: drop the private hw.aq.enable_rss_udp knob (RDTUN, default on) and add aq_rss_hashconfig(), which under options RSS returns rss_gethashconfig() and otherwise the same UDP-off default. UDP 4-tuple hashing scatters a fragmented datagram's pieces across queues because only the first fragment carries the L4 ports, so it is now off by default and re-enabled the standard way, via net.inet.rss.udp_4tuple, matching ix/ixl/mlx5. On Atlantic 1 the UDP-off action stays the existing L3L4 flow-filter workaround; only its policy source changes. TX traffic-class helper: factor the active-TC count (one per active 8-ring group, capped at HW_ATL_B0_TCS_MAX) out of aq_hw_qos_set() into aq_hw_active_tcs(), so there is a single definition of the policy; the Atlantic 2 RSS redirection table reuses it. Reviewed by: adrian Differential Revision: https://reviews.freebsd.org/D58137
aq(4): harden the attach, detach, and reset error paths
Correct several attach/detach/reset paths that either swallowed failures
or acted on undefined state.
MSI-X attach-failure double-free: aq_if_msix_intr_assign() freed the
per-RX-ring interrupts in its failure path and then returned an error, so
iflib's IFDI_DETACH freed the same irq structures again --
bus_teardown_intr() on a dangling tag and bus_release_resource() on an
already-released IRQ, panicking a box that should have simply failed to
attach. Let iflib own the teardown; drop the failure-path loop and the
now-dead index bookkeeping.
Detach loop bound: aq_if_detach() freed the per-ring interrupts looping
to isc_nrxqsets while indexing rx_rings[], which is sized by
rx_rings_count; index by rx_rings_count to match every other RX-ring
loop.
AQ_HW_WAIT_FOR final poll: the macro derived its result from the loop
counter rather than the condition, so a condition that became true on the
last iteration reported ETIMEDOUT. Worst for the acquire-on-read
firmware RAM semaphore, which was acquired in hardware but reported as a
timeout. Return based on the last evaluation of the condition.
RBL MAC reset SPI cleanup: mac_soft_reset_rbl() fired the global reset
without first tearing down the SPI/flash interface, so a flash burst in
flight left the SPI bus wedged, the RBL could not re-read flash, and the
reset returned EBUSY -- fatal at attach ("MAC reset failed: 16"). Set
bit 4 of the SPI control register (0x53c) before the global reset, as the
sibling FLB path and the Linux driver do.
Reset failure propagation: aq_hw_reset() discarded fw_ops->reset()'s
return, so a failed attach-time fw2x capability read left fw_caps == 0
permanently and stats silently froze. Propagate the error so the reset
fails and is retried.
aq_hw_init failure propagation: aq_hw_init() discarded
aq_hw_init_tx_path()/aq_hw_init_rx_path() returns and reported success,
bringing the interface up half-initialized; capture both and goto
err_exit (mainly the Atlantic 2 RX action-resolver path, which returns
EBUSY on ART semaphore timeout).
Link-state outputs: aq_hw_get_link_state() left *link_speed and *fc_neg
unwritten on early-return paths, and the caller acts on them
uninitialized, so a transient firmware get_mode() failure could fabricate
a phantom link-up at a garbage speed and program a garbage RX-pause bit.
Initialize both to safe link-down values before calling get_mode().
Reviewed by: adrian
Differential Revision: https://reviews.freebsd.org/D58138
aq(4): harden the interrupt and MAC-statistics paths
Firmware-statistics accounting and interrupt-routing fixes.
Stats delta underflow: guard the MAC statistics delta accumulation
against counter wrap or a firmware counter reset, so a snapshot smaller
than the previous one does not underflow into a huge spurious delta.
Skip stats on a failed read: aq_update_hw_stats() ignored
aq_hw_mpi_read_stats()'s return and committed the on-stack mbox into
last_stats unconditionally. On a failed read that snapshot is garbage or
zero and poisons the delta baseline (a zeroed snapshot wipes last_stats,
so the next good read double-counts). Check the return and skip the
accumulation and the last_stats commit on failure.
Mailbox/stats separation: struct aq_hw_stats served both as the raw fw1x
MCP mailbox layout and as the driver's canonical stats snapshot, so any
field added to it would silently shift the fw1x mailbox read. Give the
fw1x mailbox its own raw layout in struct aq_hw_fw_mbox and let
aq_hw_stats become purely driver-owned; with the coupling gone, add
first-class aggregate octet fields (brc/btc) that Atlantic 2 B0 firmware
can populate directly. No A1 behavior change. The raw block is a named
struct (aq_fw1x_mbox_stats) with a _Static_assert tying its size to
aq_hw_stats' matching prefix, so the fw1x memcpy cannot silently misalign
if either field list drifts. Also drop the unused FW1X_MPI_STATE_ADR /
FW1X_MPI_CONTROL_ADR macros and the redundant fw1x_get_stats() dpc
assignment that the caller immediately overwrites.
Per-speed interrupt moderation: aq_hw_interrupt_moderation_set()
hardcoded speed_index = 0, so every link speed got the 10G timer pair and
the other rows were dead. Record the negotiated rate and index the
tables by ffs(speed) - 1, reordering the rows to match the
enum aq_fw_link_speed bit positions so the index cannot drift from the
enum. Rename the two per-speed timer tables (AQ_HW_NIC_timers_table_
{rx,tx}_ -> aq_itr_timers_{rx,tx}), function-local static arrays whose
SCREAMING_CASE vendor names read like macros.
Hardware error interrupts: route both hardware error causes (interrupt
map register 0) to the admin vector so they are actually delivered.
Reviewed by: adrian
Differential Revision: https://reviews.freebsd.org/D58139
aq(4): remove dead code and tidy macros, diagnostics, and naming Non-functional cleanup, with two diagnostic corrections. Dead code: delete leftover commented-out AQ_DBG_ENTER/EXIT/PRINT calls (aq_hw.c, aq_fw2x.c, aq_irq.c, aq_main.c), a commented-out aq_nic_cfg local, the stale old-signature parameter blocks between the ring-init declarations and their bodies (aq_ring.c), a trailing note on a live statement, and the unused DumpHex() vendor debug helper (no callers; its body only compiled under AQ_CFG_DEBUG_LVL > 3). Register-write macros: parenthesize AQ_WRITE_REG_BIT's msk/shift/value arguments so a compound argument cannot mis-bind, give AQ_HW_FLUSH() an explicit hw parameter instead of capturing it from caller scope, and drop the duplicate lowercase aq_hw_write_reg[_bit] aliases (converting the 43 call sites to the uppercase spelling) so there is a single form. Diagnostics: the aq_log* family expanded through the base log macro, which ignored its level and printed unconditionally, while the error traces gated on a debug level that defaulted below LOG_ERR and so were suppressed -- backwards. Gate the base log macro the way the trace one does and default the level to lvl_error, so the once-per-event firmware reset / capability errors are visible by default while the verbose info/dump output stays opt-in. Naming: rename identifiers carried verbatim from the vendor import that do not match style -- names mixing an ALL-CAPS macro-style prefix with a lowercase tail, and a trailing underscore the vendor used as a "file-local" marker in place of static. - dbg_level_ / dbg_categories_ -> aq_dbg_level / aq_dbg_categories: these are real globals (the log/trace macros reference them from every translation unit), so the trailing underscore was never a stand-in for static; give them the aq_ namespace so the driver stops exporting generically-named global symbols. - log_base_ / trace_base_ -> aq_log_base / aq_trace_base: the internal macros behind the aq_log*/trace* families. - bootExitCode / flbStatus -> boot_exit_code / flb_status (aq_fw.c); flb_status now matches the identically-purposed variable already spelled that way in the sibling FLB-reset path. Cosmetic: terminate the ring/HW-init, MSI-X admin-handler, and media-change error messages with a newline so they are not garbled into adjacent dmesg output, and label the per-queue rx_bytes sysctl "RX Octets" (it was copy-pasted "TX Octets"). Reviewed by: adrian Differential Revision: https://reviews.freebsd.org/D58140
aq(4): add Atlantic 2 (AQC113) device support
Add support for the Marvell Atlantic 2 (AQC113/114/115/116) controllers,
a new chip generation that is not register-compatible with the Atlantic 1
parts aq(4) supports today. Adapted from the OpenBSD/NetBSD if_aq driver.
Register and device definitions (aq2_hw.h): the firmware handshake
(MIF_BOOT / MCP_HOST_REQ_INT / MIF_HOST_FINISHED), the 0x12000/0x13000
firmware interface windows, and the action-resolver table (ART) that
replaces Atlantic 1's discrete RX filters, plus the Atlantic 2 PCI device
ids and the aq_is_atlantic2() helper. Reserve a chip-feature bit
(AQ_HW_CHIP_ATLANTIC2) and add the aq_hw fields the firmware fills at boot
(ART base index, statistics interface version A0/B0). The per-VLAN-filter
resolver-tag field comes from the Linux driver; the BSD sources never
write it.
Firmware operations (aq_fwa2.c): Atlantic 2 talks to the management CPU
through the 0x12000/0x13000 register windows plus the boot handshake,
rather than Atlantic 1's mailbox in shared RAM. Implement that as a third
aq_firmware_ops vtable (reset, set_mode, get_mode, get_mac_addr,
get_stats); aq_fwa2_reboot() boots the firmware, selects the A2 ops, and
reads the version and ART base index, failing fast on the
crash-init / boot-failed bits. fwa2_set_mode advertises full duplex only
(the media model exposes no half-duplex types) and writes and acks the
link options before raising ACTIVE mode, so a forced media change does not
begin negotiation with a stale rate mask. enum aq_fw_link_speed gains
aq_fw_10M, which Atlantic 2 supports and Atlantic 1 does not.
Probe and attach: list the device ids with their media types and link
speeds (all copper; AQC113* up to 10G, AQC116C to 1G), populate
hw->device_id, and tag the generation with AQ_HW_CHIP_ATLANTIC2 so
IS_CHIP_FEATURE() recognises it uniformly. Branch firmware bring-up and
reset on the generation: aq_hw_init_ucp() and aq_hw_reset() reboot the MCP
instead of the Atlantic 1 RBL/FLB reset -- without a real datapath reset
every stop/init cycle reprograms the rings on a live, desynced RX DMA
engine and the receive path stays dead. aq_hw_init() programs the
Atlantic 2 launch-time clock ratio in place of the Atlantic 1
MRRS / TX-DMA request-limit clamp. Add an AQ_LINK_10M capability bit
(Atlantic 2 links at 10M, Atlantic 1 cannot), offer 10baseT media, and map
IFM_10_T to aq_fw_10M.
With every supported media type now present, replace the per-speed switch
statements in aq_media.c with a single {link bit, fw rate, IFM_* subtype,
Mbit/s} table -- one source of truth for the supported link speeds.
With this an Atlantic 2 card probes, brings up its firmware, reads its
MAC, and negotiates link; the RX action-resolver datapath comes next.
Reviewed by: adrian
Differential Revision: https://reviews.freebsd.org/D58141
aq(4): program the Atlantic 2 multiqueue datapath
Wire up the Atlantic 2 receive datapath: the action-resolver table (ART),
multiqueue RSS, QoS, and interrupt moderation.
RX action-resolver table: Atlantic 2 replaces Atlantic 1's discrete RX
filter registers with an ART -- hardware computes a per-packet
classification tag, then walks {tag, mask, action} rows to drop, assign a
queue, or assign a TC. aq_hw_art_filter_set() installs one row under the
ART semaphore; aq_hw_init_rx_path() enables the resolver, tags L2
unicast/broadcast, installs the unicast/all-multicast and VLAN drop rows,
and assigns every 802.1p priority to TC 0 (mirroring the Atlantic 1
user-priority map, since our RX side is a single 8-ring group in TC 0).
Tag every enabled VLAN filter in the per-filter resolver-tag field -- a
register the BSD ports never write -- because the VLAN drop row matches
resolver tag 0, so without it all tagged receive was dead under VLAN
filtering. Promiscuous mode disables the drop rows rather than toggling
the Atlantic 1 promiscuous bits; all ART callers surface a semaphore
timeout consistently. The Atlantic 1 RX_TCP_RSS_HASH and TPO2
programming is gated to Atlantic 1.
Multiqueue RSS and QoS: fill Atlantic 2's own per-TC redirection table
(AQ2_RPF_RSS_REDIR), skipping the Atlantic 1 table and its write-enable
handshake. Program Atlantic 2's smaller packet-buffer sizes, its wider
data-TC credit/weight fields, and its ring-to-TC map, using
aq_hw_active_tcs() for the TC loops.
RSS hash types: the Atlantic 2 resolver has per-protocol hash-type enables
in REDIR2, so build the mask from aq_rss_hashconfig() instead of
hardcoding every protocol -- UDP 4-tuple hashing now follows the kernel
policy (off by default) with no L3L4 flow-filter workaround, and
aq_hw_udp_rss_enable() is skipped on Atlantic 2. The kernel-to-hardware
hash-type mapping is a small static lookup table rather than a nine-branch
chain, since the two bit spaces do not share a simple shift.
Tx interrupt moderation: Atlantic 2's per-ring Tx moderation control
register lives at a different address, but its field layout matches the
value the driver already builds, so write that value straight to it; Rx
moderation is shared.
HW-validated on AQC107 <-> AQC113C: TCP RSS spreads across 7/8 RX queues
under 16 parallel flows, rx_err=0.
Reviewed by: adrian
Differential Revision: https://reviews.freebsd.org/D58142
aq(4): correct Atlantic 2 register access Four Atlantic 2 register-access corrections found in bring-up. B0 aggregate octet counters: the B0 firmware statistics interface reports only aggregate rx/tx good octets, not the per-cast breakdown A0 and Atlantic 1 provide, so every octet sysctl read a permanent zero while frame counters advanced. Populate the aggregate octet fields from the B0 buffer; aq_update_hw_stats() accumulates them directly when the per-cast octets are absent. Drop the duplicate attach-time MCP reboot: aq_hw_mpi_create() already reboots the A2 firmware to read its version and caps, then aq_hw_reset() immediately rebooted it again -- a full MCP restart plus several transaction-id-bracketed window reads, adding attach latency and a duplicate banner. Give aq_hw_reset() a reboot flag and pass reboot=false for A2 at attach; the load-bearing down/stop reboot (which resyncs A2 RX DMA across ifconfig down/up) keeps reboot=true. Skip Atlantic 1 register accesses on Atlantic 2: gate out the 0x7040 Atlantic 1 TPO write (which A2 lacks; already a no-op via the unset TPO2 feature, but Linux hw_atl2 omits it), and guard the aq_hw_mpi_read_stats() direct reads of reg_rx_dma_stat_counter7 (dpc) and the LRO counter (cprc) with !ATLANTIC2 -- those are Atlantic 1 codegen offsets that on Atlantic 2 land on unrelated registers and can report bogus input-drop / LRO counts. HW-validated on AQC107 <-> AQC113C: A1 stats unchanged, A2 IQDROPS stays 0, attach consumes one MCP reboot instead of two, bidirectional iperf3 clean. Reviewed by: adrian Differential Revision: https://reviews.freebsd.org/D58143
aq(4): observability controls and sysctl/header hygiene Fold the driver's observability and infrastructure work. Make aq_device.h self-contained: it declares struct aq_dev in terms of iflib, bitstring, socket, and ethernet types but included none of the headers that define them, compiling only because every includer happened to pull those first. Include what it uses. No functional change. Make the debug controls per-instance. The debug and debug_categories sysctls were registered per device but pointed at file-scope globals, so writing dev.aq.1.debug also changed dev.aq.0.debug and a card could not be traced in isolation. Move the level and category mask into struct aq_dev, reach them through the aq_dev back-pointer in struct aq_hw (wired up in attach_pre before the first firmware trace and guarded against a NULL deref), emit through device_printf() so each line carries its unit, and seed initial values from per-unit device hints so attach can be traced. Expose the PHY die temperature as dev.aq.N.temperature through a new firmware get_temp op: Atlantic 1 v2 reads it through the mailbox MPI control/state toggle, Atlantic 2 from the phy_health_monitor block in the OUT window (located at 0x13620 and confirmed by its ready bit). Atlantic 1 v1 has no sensor and exposes no node. Because this is the first firmware accessor iflib does not serialise, add a per-instance mutex in struct aq_hw and take it across the v2 read-modify-write in set_mode(), get_stats(), get_mode(), and get_temp(); the v1 and Atlantic 2 paths do not need it and say so. Trace the Atlantic 2 firmware path, which previously emitted nothing at any debug level (aq2_fw.c did not even include aq_dbg.h): the boot handshake, reset policy, MAC address, and link mode set/read, using the existing dbg_init and dbg_fw categories, with the per-poll mode read at detail level. Scope the driver sysctls to a context freed at detach. They were registered on the device newbus context, which newbus tears down only after DEVICE_DETACH returns, yet iflib frees the rings and softc inside DEVICE_DETACH -- a sysctl read racing detach could touch freed memory. Give the driver its own sysctl_ctx_list and free it at the start of aq_if_detach, draining in-flight readers first. Signed-off-by: Nick Price <nick@spun.io> Reviewed by: adrian Differential Revision: https://reviews.freebsd.org/D58434
aq(4): PHY thermal-shutdown handling and correctness fixes Fold the thermal-protection work and the correctness fixes that landed alongside it. Report and auto-recover from PHY thermal shutdown. The Atlantic PHYs can autonomously shut down on over-temperature, latching global fault 0x8007 and dropping the link; Atlantic 2 ships this armed, Atlantic 1 disabled. Arm it on Atlantic 1 at interface init (1E.C478.A via the MAC's MDIO controller), and recover from a trip automatically: the admin-status poll detects the fault, logs the shutdown limit and measured temperature, and holds the link down until the PHY cools, then restores it -- Atlantic 1 needs a PHY reset (1E.2681.0) with the MAC firmware running plus a full re-init, Atlantic 2 recovers on the re-init alone. New firmware ops get_phy_fault, phy_reset, thermal_arm, and get_thermal_limit back the state machine in aq_if_update_admin_status(). Make that Atlantic 1 thermal MDIO path address-correct and fail-safe. The direct-MDIO helpers hardcoded the Clause-45 port address to 0, but it is strap-selectable: on a board whose PHY answers elsewhere every thermal op targeted nothing, so arming silently no-oped and the post-trip reset never cleared the latch. Discover the address by scanning ports 0..31 for a PMA/PMD identifier and form it as (phy_id << 5) | mmd, marking it valid only when a PHY actually answers. aq_fw2x_phy_read also returned 0 on a semaphore timeout, indistinguishable from a real 1E.C478 == 0, so thermal_arm could zero live provisioning bits; give the read an error return and gate thermal_arm and get_thermal_limit on it. Bound the multicast filter slot index. aq_mc_filter_apply() programmed slot count + 1 and bailed only at count == AQ_HW_MAC_MAX (33), one address too late, so a 33rd entry raced in between the if_llmaddr_count() snapshot and the if_foreach_llmaddr() walk drove an out-of-bounds MMIO write to slot 33. Fire the guard at AQ_HW_MAC_MAX - 1, and also reject index >= AQ_HW_MAC_MAX in aq_hw_mac_addr_set() where the slot becomes an RPF register offset. Correctness and safety fixes: initialize the sysctl context in attach_pre so the iflib fail-path detach cannot sysctl_ctx_free() an uninitialized list (a page fault when MSI/MSI-X is denied); range-check the Atlantic 2 action-resolver table index, taken verbatim from a firmware-supplied base, before writing the ART registers; and accumulate statistics deltas as unsigned, since AQ_SDELTA discarded a forward delta of 2^31 or more at 10G across a stretched admin poll. Signed-off-by: Nick Price <nick@spun.io> Reviewed by: adrian Differential Revision: https://reviews.freebsd.org/D58435
aq(4): clean up diagnostics and remove dead code
Non-functional cleanup, no change in behavior.
device_printf() already prefixes each line with the device name, so the
inline "atlantic:" token in the status and error messages produced a
doubled prefix and diverged from the trace macros; remove it so all
output carries one uniform "aqN:" prefix. Compile the RX/TX descriptor
tracers only when AQ_CFG_DEBUG_LVL > 2 and make them no-op macros
otherwise, so the default build no longer pays a cross-TU call plus
argument evaluation per descriptor.
Drop enum aq_dev_state, struct aq_rx_filters, and struct aq_vlan_tag,
which have no remaining references now that VLAN state lives in a
bitstr_t. Replace the four identical aq_sysctl_print_{tx,rx}_{head,tail}
handlers, each carrying a dead write path on a read-only oid, with one
aq_sysctl_print_ring_ptr that selects the accessor from arg2. Reduce the
thermal and PHY-recovery comments to single terse lines that keep the
load-bearing register numbers and the A1-vs-A2 recovery difference.
Signed-off-by: Nick Price <nick@spun.io>
Reviewed by: adrian
Differential Revision: https://reviews.freebsd.org/D58436
aq(4): mailbox, flow-control and firmware error-handling fixes Fold the whole-driver-review correctness and hardening fixes for the firmware and hardware layers. Advance the firmware-mailbox address per word in aq_hw_fw_downld_dwords(): on B1 silicon each loop iteration waits for the mailbox address register to differ from the expected address, but it was set once and never moved, so after the first word every wait returned immediately and read stale data. Advance it four bytes per word. B0 is unaffected (it polls the busy bit). The same function also left err set to ETIMEDOUT after successfully force-recovering the RAM CPU semaphore; the transfer loop is guarded by "--cnt && !err", so it ran zero iterations and returned a timeout with an untouched buffer, making the recovery path dead code. aq_hw_get_mac_permanent() ignored the get_mac_addr() error and then examined a buffer the firmware op never wrote on failure. A fresh softc is zero, so the "invalid address" test fired, a random locally administered MAC was substituted, and err was overwritten with 0 -- a transient mailbox failure produced a card that attached with a different MAC every boot. Fail instead; the random-address fallback still covers a genuinely blank or multicast burned-in address. aq_fw1x_reset() discarded the same download's return value and then read transaction_id out of an uninitialized stack struct, so propagate that error too. Encode RX-only flow control as PAUSE|ASYM_PAUSE rather than PAUSE alone: firmware 2.x/3.x has no independent RX-only bit, so the old encoding advertised symmetric pause when RX-only was requested. The MPI_INIT path also never cleared the pause bits before OR-ing in the requested ones, so flow control could be enabled and never disabled; clear them first, as the Atlantic 2 and Linux implementations do. Reject single-vector MSI in aq_if_attach_post() the same way legacy INTx is rejected: ift_legacy_intr is NULL, so no driver filter would acknowledge the not-clear-on-read, auto-masked device interrupt status; every supported Atlantic device provides MSI-X. Propagate firmware and MDIO errors instead of discarding them. The fw2x MDIO primitive returned a data word with no way to report a controller timeout; give aq_fw2x_mdio_op() a status return and a data out-parameter, propagate it through phy_write/read/reset/thermal_arm, and stop advancing the thermal recovery state machine when a PHY reset fails. Use that error to end the PHY address scan early: aq_fw2x_init_phy_id() probed all 32 MDIO ports even when the controller itself was timing out, spending up to ten seconds under fw_mtx and the iflib context lock. aq_fw2x_reset() also drove the shared MIF mailbox without fw_mtx, unlike every other fw2x mailbox user, so it could interleave with the temperature sysctl and load the capability mask from the wrong window. aq_hw_mpi_set() can return ETIMEDOUT when the Atlantic 2 shared firmware buffer is not acknowledged; aq_hw_init() now aborts through its error path rather than enabling rings with an unaccepted link state, and aq_if_init() logs the later link-speed error. Retry a failed initialization instead of leaving the link down. ifdi_init has no return value, so iflib marks the interface running once aq_if_init() returns; a propagated firmware-ack failure would otherwise leave it running with no initialized hardware and no recovery. Record the failure and retry from the admin task via iflib_request_reset(), paced by the once-per-second timer, giving up after a bounded number of attempts. Ring and queue start failures are deliberately left to the existing diagnostic, since they leave the remaining queues usable. Reviewed by: adrian Differential Revision: https://reviews.freebsd.org/D58437
aq(4): interface lifecycle and link-state fixes aq_if_init() programmed the address captured at attach, so an address set with "ifconfig ether" or by lagg(4) enslavement was never written to unicast filter slot 0: the interface transmitted with the new address but the MAC still filtered on the old one, so it received nothing. Copy the current if_getlladdr() the way the other iflib drivers do. The link state could latch UP forever. aq_if_stop() cleared linkup before calling aq_if_update_admin_status(), which suppressed the LINK_STATE_DOWN transition the "link was UP" branch would have made. Announce the down transition directly from aq_if_stop() instead, and do not poll the admin status there at all: the MAC has just been reset, so a stale link reading would re-announce the link as up. The admin task itself had to stop reporting a link on a stopped interface. iflib runs it while either IFF_DRV_RUNNING or IFF_DRV_OACTIVE is set, and iflib_stop() sets OACTIVE, so the task kept polling after the stop and re-announced LINK_STATE_UP behind the driver's back. Treat a non-running interface as having no link. A lagg(4) parent otherwise keeps hashing flows onto a port whose carrier is gone, because LAGG_PORTACTIVE tests if_link_state together with IFF_UP. Stop the rest of the task there as well: the PHY thermal poll and the initialization retry both end in iflib_request_reset(), and _task_fn_admin() acts on that with no test of its own, so either could re-initialize an interface the operator had just taken down. aq_if_update_admin_status() also only reacted to transitions in and out of zero speed, so an autoneg downshift that kept the link up left if_baudrate, ifmedia, RX pause and interrupt moderation programmed for the old speed. Track the announced speed and re-run that work when it changes. aq_if_suspend() resets the MAC and stops the rings, but iflib_device_suspend() only calls IFDI_SUSPEND and never stops the interface, leaving IFF_DRV_RUNNING set over a suspended device. Clear it. Signed-off-by: Nick Price <nick@spun.io> Reviewed by: adrian Differential Revision: https://reviews.freebsd.org/D58473
re(4): quiesce RTL8168G+ and reset before freeing buffers in re_stop() The STOPREQ command written by re_stop() is not defined for RTL8168G and later; issuing it can wedge the MAC. Replace it on those parts with the vendor-documented sequence: * settle delay * bounded poll for Tx queue empty * clear TE/RE * then bounded poll of the MCU command register (0xD3) FIFO-empty bits. Also reset the controller before the Rx/Tx buffer free: a controller that has not quiesced keeps DMAing stale, still-owned descriptors pointing at freed mbufs (use-after-free under INVARIANTS, cross-NIC mbuf corruption reported in the PR). Adds the RL_MCU_* register definitions. All waits are bounded; error paths only. * iperf3 --bidir at line rate against RTL8168H (XID 0x541); previously wedged the controller until power cycle, with the quiesce the reset path recovers. * Deployed in production on an RTL8168H fleet since 2026-07-01. Reviewed by: adrian Differential Revision: https://reviews.freebsd.org/D58276 PR: kern/https://bugs.freebsd.org/bugzilla/show_bug.cgi?id=166724
re(4): re-arm the Tx doorbell when re_txeof() leaves a non-empty ring On PCIe parts a TxPoll request can be lost when packets are queued in quick succession, leaving owned descriptors with no transfer in progress until the watchdog fires. re_txeof() runs from the interrupt handlers, re_tick() and re_watchdog(), so re-writing TXSTART whenever the ring is still non-empty turns a potential 5-second stall into at most one tick. One register write on a path that already took an interrupt; fast path untouched. * Sustained bidirectional load on RTL8168H; no Tx stalls, no throughput regression at 941 Mbps line rate. Reviewed by: adrian Differential Revision: https://reviews.freebsd.org/D58277 PR: kern/https://bugs.freebsd.org/bugzilla/show_bug.cgi?id=166724
re(4): recover Tx completions whose MSI was swallowed in re_intr_msi() A Tx completion that raises a status bit between the ISR ack at the top of re_intr_msi() and the IMR re-enable at the bottom is never re-signalled: these controllers do not re-assert MSI for an already-set status bit (this is why hw.re.msi_disable is a known workaround in the PR). Re-read ISR before re-enabling; if a Tx bit is pending, ack just that bit, reap the ring and restart the queue. Rx bits are deliberately left set so they re-arm the interrupt normally and Rx moderation state is untouched. Also flush the posted IMR write. Mirrors what the INTx path already achieves via the loop in re_intr(). * MSI interrupt mode on RTL8168H under load; "missed Tx interrupts" watchdog recoveries no longer occur. Reviewed by: adrian Differential Revision: https://reviews.freebsd.org/D58278 PR: kern/https://bugs.freebsd.org/bugzilla/show_bug.cgi?id=166724
re(4): harden re_watchdog() recovery and log controller state Distinguish the two failure classes from the PR in a single log line (ring indices, ISR/IMR, TXCFG, interrupt mode): lost interrupt vs genuine DMA stall. Bail out instead of re-initializing when the controller reads back all-ones (fallen off the bus; reinit cannot help). Re-assert the driver's existing ASPM-disabled policy before reinit, since firmware/power transitions re-arming L0s/L1 is a documented stall trigger. Diagnostics-only on the recovered path; no fast-path change. * Field diagnostics running on an RTL8168H production fleet; the log format distinguishes lost-doorbell / DMA-stall / dead-controller without a debug build. Reviewed by: adrian Differential Revision: https://reviews.freebsd.org/D58279 PR: kern/https://bugs.freebsd.org/bugzilla/show_bug.cgi?id=166724
re(4): add hw.re.aspm_disable loader tunable re(4) has unconditionally disabled ASPM L0s/L1 and CLKREQ at attach for years; on laptops this costs 200mW+ (requested by adrian@ in the PR). Make it a tunable following the existing hw.re.* pattern: * default 1 keeps today's behavior; * 0 preserves the firmware-configured ASPM state at attach and skips the watchdog re-assert from the previous revision. Documented in re.4. * Verified on RTL8168H (XID 0x541): with hw.re.aspm_disable=0, attach no longer logs "ASPM disabled" and pciconf -lcb shows the firmware Link Control state preserved -- including Clock PM, which the unconditional code previously cleared. * Default (1) is behaviorally identical to the current driver. * Note the tunable also stops the driver clearing CLKREQ, a small power win even where firmware leaves L0s/L1 off. Reviewed by: adrian Differential Revision: https://reviews.freebsd.org/D58280 PR: kern/https://bugs.freebsd.org/bugzilla/show_bug.cgi?id=166724
Provide pc_small_core for i386 too to fix an i386 build break from x86 code referring to it. It won't be set. Reviewed by: aokblast, kib Fixes: https://cgit.freebsd.org/src/commit/?id=7b26353a59d6 ("hwpstate_intel: Disable package control on hybrid CPU") Differential Revision: https://reviews.freebsd.org/D58335
bus_{read,write}_8 are macro wrappers around the corresponding bus_space
functions in sys/bus.h, so implementing bus_{read,write}_8 won't work.
Implement the underlying bus_space function instead.
Reviewed by: jrtc27, rlibby
Fixes: https://cgit.freebsd.org/src/commit/?id=9313f6b01485
Sponsored by: The FreeBSD Foundation
Differential Revision: https://reviews.freebsd.org/D58301
The accumulated count of a process-mode counting PMC is kept in a 64-bit software counter and seeded into the hardware counter at every context switch in. Hardware counters are narrower than that - each PMC class discovers and records its own counter width, e.g. 48 bits on current x86 (queried from CPUID on Intel, architectural on AMD) - so once the accumulated count approaches the end of the hardware counter range, the counter wraps during a time slice and the value read back at switch out is smaller than the value seeded. The increment was computed assuming a full 64-bit counter: on INVARIANTS kernels a long enough counting run panics with "negative increment" the moment the accumulated count first crosses the hardware counter range, and on other kernels the totals silently lose a full counter range per wrap. Compute the increment modulo the per-class hardware counter width instead, in both places that accumulate switch-out deltas. Reviewed by: adrian MFC after: 2 weeks Assisted-by: Claude Code (Fable 5) Differential Revision: https://reviews.freebsd.org/D58340
A process-mode PMC's runcount tracks how many CPUs currently have it
loaded in hardware. It is decremented only by the context-switch-out
and process-exit reclaim paths, both of which the scheduler invokes
only for processes flagged P_HWPMC. Detaching a target that still has
the PMC live in hardware dropped the target and cleared P_HWPMC without
taking the PMC off the hardware or dropping the runcount reference, so
the reference leaked. A subsequent release then spun in
pmc_wait_for_pmc_idle() forever waiting for the runcount to reach zero:
on an INVARIANTS kernel this panics ("waiting too long for pmc to be
free"), otherwise it is an unkillable loop holding the hwpmc lock. Any
process able to allocate a PMC can trigger this by attaching a counting
PMC to itself and detaching it before releasing.
Take the PMC off the hardware and drop the runcount reference as part
of detaching, before P_HWPMC is cleared: reclaim it from the detaching
thread's own CPU directly, and, when the detach removes the PMC's last
target, wait for any references held by the target's other threads to
drain while P_HWPMC is still set (they can no longer reload it).
Reviewed by: adrian
MFC after: 2 weeks
Assisted-by: Claude Code (Fable 5)
Differential Revision: https://reviews.freebsd.org/D58342
Some UVC devices (e.g. Logitech C920) expose more than 8 Processing Unit descriptors, causing "too many PU descriptors found!" errors. Increase both limits from 8 to 32 to accommodate such devices.
Import the device quirk system from OpenBSD to handle UVC devices that need special handling. This includes: - UVIDEO_FLAG_ISIGHT_STREAM_HEADER: non-standard streaming header - UVIDEO_FLAG_REATTACH: needs reattach after firmware upload - UVIDEO_FLAG_VENDOR_CLASS: incorrectly reports as vendor class - UVIDEO_FLAG_NOATTACH: device not supported - UVIDEO_FLAG_FORMAT_INDEX_IN_BMHINT: format index in bmHint Add quirks table with known devices and lookup function. Add iSight stream header decoder for Apple iSight cameras. Obtained from: OpenBSD
Sponsored by: The FreeBSD Foundation MFC after: 1 week Differential revision: https://reviews.freebsd.org/D56005
A bunch of drivers weren't properly converted. I mistakenly put a call to ieee80211_output_seqno_assign() wherever the crypto header was added, which isn't exactly correct. There are plenty of drivers which don't share enough of their raw and normal transmit path code for that to hold true. So after some manual review, it looks like I've captured the places (outside of iwn(4) which I committed earlier) where I missed ieee80211_output_seqno_assign() calls. * For bwi(4) and bwn(4) I refactored it out into a place that is common enough and happens in the same lock hold window, so it's serialised. * For the rest, it's just plain missing from the raw path. Locally tested: * ural(4) * ral(4) * bwi(4) Differential Revision: https://reviews.freebsd.org/D58098
firewire: add warn-only CRC validation for CSR ROM directories Implemented crom_crc_valid() helper to validate IEEE 1394 config ROM CRC-16 checksums. Skipped root header CRC validation since csrhdr.crc_len cover the entire ROM body which is not fully read at header parse time. Per-directory CRC checks below catch corruption where it needed. Reviewed by: adrian Differential Revision: https://reviews.freebsd.org/D58307
firewire: drain pending xfers after callout stop in detach Removes a TODO that predates the existing drain call. Reviewed by: adrian Differential Revision: https://reviews.freebsd.org/D58308
firewire: force root change when root node is not cycle master capable When a FireWire bus resets, all devices negotiate who is the new boss. when we detect the root node can't be cycle master, we send a PHY config packet that forces a reelection. Reviewed by: adrian Differential Revision: https://reviews.freebsd.org/D58309
firewire: remove dead code across the subsystem Reviewed by: adrian Differential Revision: https://reviews.freebsd.org/D58310
firewire: replace magic numbers with named constants No functional change. Reviewed by: adrian Differential Revision: https://reviews.freebsd.org/D58311
Sponsored by: Klara, Inc. Sponsored by: NetApp, Inc. MFC after: 1 week Fixes: https://cgit.freebsd.org/src/commit/?id=6d0001d44490 ("nvme: add support for DIOCGIDENT") Reviewed by: bnovkov, imp Differential Revision: https://reviews.freebsd.org/D58357
The uvideo driver freed the mmap buffer (contigmalloc'd) in several paths (VIDIOC_STREAMOFF, last close, detach) without coordinating with the lifetime of existing user-space mmap mappings. This could lead to use-after-free when user-space continued to access the mapped memory after the backing pages had been freed. Fix this by switching from the simple d_mmap callback to d_mmap_single with custom cdev_pager_ops, and by attaching the contig buffer to a single shared vm_object created at REQBUFS time: - uvideo_reqbufs() allocates a uvideo_mmap_state (independent of the softc) and a shared vm_object via cdev_pager_allocate() that spans the whole buffer; the softc holds one reference to it. - uvideo_cdev_mmap_single() simply hands out additional references to that shared object; the requested offset selects which buffer is mapped. The VM system tracks mapping lifetime through the object reference count, so no per-mapping bookkeeping is needed. - uvideo_pg_ctor/uvideo_pg_dtor validate the mapping and free the contig buffer together with the state when the last reference (softc's own or a user mapping) is dropped. - uvideo_pg_fault installs a fictitious page for the backing physical address, following the canonical device-pager pattern: update the passed-in page in place when it is already fictitious, otherwise allocate a fake page and vm_page_replace() the busy placeholder, so that dev_pager_dealloc() does not deadlock. - uvideo_vs_free_frame() drops the softc's reference instead of contigfree()'ing directly; if mappings still exist the buffer stays alive until the last uvideo_pg_dtor(). - VIDIOC_STREAMOFF no longer frees the buffer (per V4L2 spec). - Last close always releases the buffer (deferred if mappings exist). - The mmap_state outlives the softc, so the pager dtor can safely free the buffer even after device detach. Reported by: 章鱼哥 (@aipyapp) (www.aipyaipy.com) Reported by: Chris Jarrett-Davies of the OpenAI Codex Security Team
This fixes an issue with the Solo2 (and likely some of the Nitrokey
family) where hangs would occur with OpenSSH- it issues a CANCEL prior
to closing the device unconditionally, and without draining the read
endpoint we end up seeing the response to that CANCEL the next time
OpenSSH tries to connect. This throws the entire command/response
sequence out of whack.
This call used to break Yubikeys in some situations, but the fix that
landed in 28d85db46b48 ("xhci: Do not drop and add bits in xhci") seems
to have addressed that- presumably we sometimes end up stopping the
command and desyncing at the controller level. This probably implies
that we need a SYNCWRITE HID quirk, but that requires a little more work
in usbhid_sync_xfer() and this doesn't seem to cause any problems in
normal usage.
Reviewed by: aokblast, wulf
Differential Revision: https://reviews.freebsd.org/D58199
Replace bare EINVAL in AMD/IBS allocation and config-validation with EXTERROR(), so a failed pmc(3) allocation names the check and value. Register HWPMC_AMD in exterr_cat.h and the generated filenames.h. Signed-off-by: Andre Silva <andasilv@amd.com> Reviewed by: Ali Mashtizadeh <ali@mashtizadeh.com>, mhorne Sponsored by: AMD Pull Request: https://github.com/freebsd/freebsd-src/pull/2180
Annotate validation failures in the PMC syscall handlers (allocate, attach, read/write) with EXTERROR(), so pmc(3) callers see which precondition failed, not a bare errno. Register HWPMC_MOD in exterr_cat.h and the generated filenames.h. Signed-off-by: Andre Silva <andasilv@amd.com> Reviewed by: Ali Mashtizadeh <ali@mashtizadeh.com>, mhorne Sponsored by: AMD Pull Request: https://github.com/freebsd/freebsd-src/pull/2180
vtnet: Retry feature negotiation without offloads A device is permitted to reject an otherwise valid subset of its offered features by refusing to accept FEATURES_OK (VirtIO v1.3, 2.2.2). Apple's Virtualization.framework does this in practice; it treats the offered CSUM/TSO offloads as all-or-nothing, while vtnet's default request contains only part of that group because of hw.vtnet.lro_disable that would drop the guest TSO bits, thus negotiation fails and the device does not attach. If FEATURES_OK is rejected, retry the negotiation once with every offload-related feature stripped. Changing the feature set after a failed FEATURES_OK requires re-initialising from device reset (VirtIO v1.3, 3.1.1), so the retry goes through virtio_reinit(). A NIC without offloads is preferable to no NIC at all. Devices that accept the initial feature set are unaffected, while those that also reject the reduced set continue to fail attachment as before. Signed-off-by: Faraz Vahedi <kfv@kfv.io> Reviewed by: adrian Pull Request: https://github.com/freebsd/freebsd-src/pull/2322
vtnet: Implement VIRTIO_NET_F_GUEST_ANNOUNCE When the device sets VIRTIO_NET_S_ANNOUNCE in the config status field, for example after a VM migrates to a new host, announce the interface's presence on the network so peers and switches learn the new attachment point, then acknowledge the request with the VIRTIO_NET_CTRL_ANNOUNCE_ACK control command, as per VirtIO v1.3, 5.1.6.5.4. The announcement raises iflladdr_event: the stack sends gratuitous ARPs and unsolicited neighbor advertisements for the interface's addresses, and stacked interfaces such as vlan(4) propagate the event and announce theirs as well. The event handlers may sleep, so the work is deferred from the config change interrupt to a task on taskqueue_thread; that context also allows the acknowledgement to be skipped safely if the interface was stopped in the meantime, in which case the device keeps the bit set and the request is re-delivered with the next config change interrupt. Signed-off-by: Faraz Vahedi <kfv@kfv.io> Reviewed by: adrian Pull Request: https://github.com/freebsd/freebsd-src/pull/2322
vtnet: Accept VIRTIO_NET_F_CTRL_RX_EXTRA Although the driver does not issue the extra receive-mode commands accepting the feature is harmless and some devices, notably Apple's Virtualization.framework, offer their control-queue features as a group and refuse FEATURES_OK unless the whole set is acknowledged. Signed-off-by: Faraz Vahedi <kfv@kfv.io> Reviewed by: adrian Pull Request: https://github.com/freebsd/freebsd-src/pull/2322
These kernconfs were missed in the previous commit. PR: https://bugs.freebsd.org/bugzilla/show_bug.cgi?id=289236 Reviewed by: kib Fixes: https://cgit.freebsd.org/src/commit/?id=f38cbefef8090f3363e5685c5a3b30ffbf1d3ad0 MFC after: 3 days Sponsored by: The FreeBSD Foundation
uvideo: replace contigmalloc with OBJT_PHYS vm_object for mmap buffer Allocate the mmap buffer via phys_pager_allocate() and map it into kernel space with vm_map_find()/vm_map_wire(), instead of a custom cdev_pager backed by contigmalloc. phys_pager_allocate() is required over a bare vm_object_allocate(OBJT_PHYS) to initialise un_pager.phys.ops, otherwise phys_pager_getpages() NULL-derefs during vm_map_wire(). Reviewed by: markj Reported by: markj Differential Revision: https://reviews.freebsd.org/D58394
uvideo: validate frame size before mmap buffer allocation dwMaxVideoFrameSize comes from the USB probe/commit response and is not validated. reqbufs() computed buf_size_total with signed int arithmetic and no bound, so a bogus value could wrap the product to a small size and yield a too-small buffer with a huge sc_mmap_buffer_size, causing out-of-bounds writes from the USB transfer callbacks. Bound the frame size against sc_max_fbuf_size and use overflow-checked size_t arithmetic for the total and per-buffer offsets. Reported by: emaste
uvideo: lock the mmap queue and read path qbuf(), dqbuf() and read() manipulated sc_mmap_q / sc_mmap_cur / sc_frames_ready without sc_mtx, racing with the USB transfer callbacks (producer) that run under the mutex. This could corrupt the queue or trigger use-after-free. Take sc_mtx around qbuf(), use mtx_sleep() and protect the queue operations in dqbuf(), and use mtx_sleep() with a snapshot of sc_fsize in read(). Also reject S_FMT and S_PARM with EBUSY while streaming: both re-negotiate the probe/commit controls with the device, which disrupts the active USB transfers (a second client opening the device would otherwise freeze the first one's stream).
uvideo: bounds-check frame interval reads against bLength Frame interval data is read from device-supplied frame descriptors whose bLength may be shorter than the number of intervals declared by bFrameIntervalType. The continuous branch of uvideo_enum_fivals() read three intervals unconditionally, and the discrete branch checked the pointer but not the four bytes that UGETDW() reads, so a short or malformed descriptor could read past bLength and leak adjacent kernel memory to userspace. uvideo_vs_parse_desc_frame_max_rate() had the same class of off-by-up-to-three-bytes read. Compute the available bytes from bLength and validate before each read. Reported by: emaste
uvideo: track streaming ownership per-fd and free buffers on STREAMOFF The driver shared a single streaming state and buffer pool across all open file descriptors, so a second client (e.g. another browser tab) could disrupt the first: its cleanup STREAMOFF would tear down the active stream, and stale buffers prevented re-acquisition. Add per-fd state via devfs cdevpriv tracking whether this fd started streaming. STREAMOFF and close from a non-streaming fd are no-ops. STREAMOFF from the streaming fd stops the stream and frees the buffers so that a new fd can re-acquire the camera. DQBUF returns EPIPE immediately when buffers are freed instead of waiting for a timeout.
uvideo: fix close/detach race on streaming teardown detach() stopped streaming and called uvideo_vs_close() before destroy_dev(), so a concurrent close() could race the teardown and call uvideo_vs_close() a second time (double usbd_transfer_unsetup), and mtx_destroy() could race a close still holding sc_mtx. sc_streaming was also read without the lock in both paths. Reorder detach() to call destroy_dev() first so all in-flight cdev methods drain before any teardown. Read sc_streaming under sc_mtx in both detach() and the last-close safety net.
uvideo: use size_t for sc_mmap_count and loop index in reqbufs
uvideo: validate frame descriptors and fix integer overflows in size computation
uvideo: fix printf type Reported by: vishwin
Reviewed by: aokblast, kib, olce Differential Revision: https://reviews.freebsd.org/D58336
Some UAC2 devices expose a single Clock Source entity that is shared between their playback and capture interfaces (it appears in both the output and input clock bitmaps). On such a device uaudio(4) programs the sample rate for both directions when a stream starts. If playback runs at a 44.1 kHz-family rate while the idle capture channel is left at its 48 kHz-family default, the capture SET_CUR(UA20_CS_SAM_FREQ_CONTROL) is issued after the playback one and overwrites the rate on the shared clock. The device then runs at ~48 kHz while the playback stream carries 44.1 kHz data. Consuming samples faster than they arrive, the device repeatedly runs out of data, loses sync with the playback stream, and re-locks onto it (audible dropouts, front-panel play/idle flicker). The 48 kHz family is unaffected because both directions then agree on the rate. Fix it in three parts: - Add a shared-clock guard: before issuing SET_CUR to a clock id, if that clock is shared between playback and capture and the other direction is already streaming at a different rate, skip it. The first active stream owns the clock; a later one follows it. - When the recording channel is auto-started only as a source of jitter information for asynchronous playback, align its nominal rate to the playback rate before starting it, so it neither reprograms the shared clock to a conflicting rate nor produces mismatched frame sizes. - Always submit the explicit-feedback SYNC transfer so dev.pcm.%d.feedback_rate stays live as a diagnostic even when a capture stream is present. Reproduced on an OKTO RESEARCH DAC8 STEREO (0x152a:0x88c5), whose vestigial capture interface never streams; the same device plays the 44.1 kHz family correctly under Linux's snd-usb-audio. As a side effect, this patch also fixes the sample rate bug mentioned in the BUGS section of sound(4)'s man page, where a device needs to have the same sample rate set for both playback and recording in order to work properly. PR: https://bugs.freebsd.org/bugzilla/show_bug.cgi?id=295933 Assisted-By: Claude Opus 4.8 (claude-opus-4-8) Signed-off-by: giacomo <delleceste@gmail.com> MFC after: 2 weeks Reviewed by: christos Pull-Request: https://github.com/freebsd/freebsd-src/pull/2323
The fixed 128 KiB secondary buffer cap dates from stereo-sized streams. High channel-count or high sample-width OSS streams can consume most of that budget in one graph quantum, leaving too little room for capture catch-up or playback headroom. Keep 128 KiB as the low-rate floor, but derive the effective soft-ring cap from the channel byte rate, clamped to 4 MiB. Use that per-channel cap when resizing the soft buffer and when clamping SNDCTL_DSP_SETFRAGMENT requests. Also clamp SNDCTL_DSP_LOW_WATER to the current soft-buffer size so an impossible readiness threshold cannot make poll/select wait forever. MFC after: 3 weeks Reviewed by: christos Differential Revision: https://reviews.freebsd.org/D58064
Prevent infinite loop in uvideo_vs_negotiation() when a USB camera reports step=0 in its continuous frame interval descriptor. Cast fbuf_size calculation to uint64_t to avoid int overflow for large width/height/bpp combinations. Reported by: emaste
Don't coerce errors to EINVAL, which isn't correct for mtx_sleep's failure cases. Sponsored by: The FreeBSD Foundation
Our USB TRB buildup subroutines were previously difficult to follow. In setup_generic_chain_sub(), the routine filled TRB packets based on the characteristics passed by the caller and the current state (for example, whether the TRB was the last in the TD). However, most TRB types (except Normal TRBs) cannot be shared across TDs. To simplify the logic, refactor xhci_setup_generic() so that TRBs are constructed according to their transfer type, with dedicated helper functions for each TRB type. Sponsored by: The FreeBSD Foundation Assisted-by: Claude Code (Opus 4.6, Opus 4.8(1M) and Sonet 5.0) Differential Revision: https://reviews.freebsd.org/D57130
When a USB HID device triggers identify, the grandparent is usbhid on a USB hub. Calling iicbus_get_addr() on a non-iicbus device hits a KASSERT panic. Reviewed by: adrian Differential Revision: https://reviews.freebsd.org/D58432
Currently, USB request not distinguished different error and always return EIO. However, some error are recoverable or ignorable in userspace. Therefore, we preserve the meaning of different error to userspace then allow userspace to decide how to use the return error. Reviewed by: adrian Sponsored by: The FreeBSD Foundation Differential Revision: https://reviews.freebsd.org/D52244
Reviewed by: markj MFC after: 2 weeks Sponsored by: The FreeBSD Foundation Differential Revision: https://reviews.freebsd.org/D56311
e1000: Defer link-up notification until after TSO reset em_automask_tso() changes the enabled TSO capabilities when the link moves between 10/100 and 1000 Mb/s. A running interface must be reinitialized to apply the new capability set. Do not publish LINK_STATE_UP until the requested iflib reset has completed. Replace link_active with an explicit state machine that distinguishes the physical link, its publication to iflib, and an outstanding reset barrier. Preserve that barrier across a link flap with DOWN_RESET_PENDING, and only publish DOWN if UP was previously published. Only request a reset for a running interface or for an initialization while the interface is administratively up. In other states the next initialization will apply the capability changes, avoiding a reset request that iflib's admin task could discard. Reviewed by: Faraz Vahedi <kfv@kfv.io> Fixes: https://cgit.freebsd.org/src/commit/?id=2ddf24f8f525 ("e1000: Automask TSO on lem(4)/em(4) 10/100 Ethernet") MFC after: 1 week
e1000: fix 82574 MSI-X interrupt throttling em_newitr() and the per-queue interrupt_rate sysctl both tested que->msix to decide whether an 82574 is running in MSI-X mode. 0 is a valid MSI-X vector so queue 0 was misclassified as legacy/MSI. Test sc->intr_type == IFLIB_INTR_MSIX instead. While here, index the tx EITR read by tque->msix rather than tque->me so it matches the register em_newitr() actually writes; the two differ once tx_num_queues exceeds rx_num_queues. Also seed que->itr_setting in em_initialize_receive_unit() with the rate the hardware was just programmed with. Otherwise an itr_setting left over from AIM across an interface re-init makes the change detection in em_newitr() suppress the write that would restore it, leaving the hardware at the default rate while software believes otherwise. Fixes: https://cgit.freebsd.org/src/commit/?id=3e501ef89667 ("e1000: Re-add AIM") MFC after: 3 days
e1000: fix rx accounting for multi-descriptor packets The receive paths accumulate ri->iri_len across the descriptors making up a packet, then add that running total to rxr->rx_bytes on every iteration of the loop. A packet spanning descriptors of length l1, l2 and l3 thus contributes 3*l1 + 2*l2 + l3 instead of l1 + l2 + l3. Single descriptor packets, the common case, are accounted correctly, so this only shows up on jumbo frames. Add the per descriptor length instead. iflib memsets the if_rxd_info before each isc_rxd_pkt_get() call, so summing len gives the same total as the final iri_len, and the frame error path that returns without incrementing rx_packets keeps counting bytes exactly as before. MFC after: 1 week
e1000: make AIM counter sampling coherent Sample free-running counters by delta instead of clearing them from the interrupt filter, which can race their producers. Publish byte and packet counts together at the TX and RX doorbells so each sample is coherent. Aggregate every TX ring assigned to the interrupt vector so unequal RX and TX queue counts are safe. Count RX bytes only after a frame is accepted. MFC after: 1 week
e1000: synchronize interrupt moderation state Keep the saved EITR and PBA values synchronized with hardware across reinitialization. Correct EITR encoding, decoding, and MSI-X register selection, and reject nonpositive fallback rates. Treat only sub-gigabit links as sub-gigabit and apply the packet-buffer fallback without permanently disabling AIM. MFC after: 1 week
e1000: restore packet-size AIM Restore the packet-size calculation introduced in a69ed8dfb381 and used by igb(4) until the iflib conversion in f2d6ace4a684. It derives interrupt holdoff from average packet size, so RSS queue count does not change its behavior. The calculation follows the pre-iflib code. Retain the current normal and low-latency rate caps, and keep the current setting when an interval has no usable sample. Fixes: https://cgit.freebsd.org/src/commit/?id=3e501ef89667 ("e1000: Re-add AIM") MFC after: 1 week
e1000: count TSO wire segments in the AIM counters The transmit paths billed one packet of ipi_len bytes per request. For TSO that is the whole unsegmented payload, up to 64KB, so the average size the moderation calculation sees is not a size that appears on the wire. Count the segments the hardware will put on the wire and the header each of them carries. Non-TSO accounting is unchanged. MFC after: 1 week
e1000: Retry transient MDIC failures on modern PCH Some Meteor Lake and newer systems sporadically fail an MDIC PHY transaction while the MAC and PHY clocks synchronize. Retry twice before reporting the transaction failure. Disable retries around PHY interface transitions where an MDI error is expected. Preserve and restore the configured retry count on every exit from those flows. This follows DPDK commit bdca22d62ff0, extended to the PTP and NVP PCH types. MFC after: 2 weeks
e1000: Reconfigure modern PCH K1 clock synchronization Meteor Lake and newer PCH generations can lose packets while the MAC and PHY clocks synchronize. Move K1 power-down to P1 and extend the PHY K1 exit timeout before PHY access and after reset. Use the longer 1 Gb/s PLL clock-gate timeout added by Linux so K1 can remain enabled without the power penalty of disabling it. Apply the workaround through the newer PTP and NVP generations. This follows DPDK commits ba54bdc79d94 and d88ef2356ecc, with the longer exit time observed in Linux 578294b8b60d. MFC after: 2 weeks
e1000: Compare decoded PCH LTR latencies The LTR encoding combines a value and a nonlinear scale, so encoded values cannot be compared directly. Decode both the device latency and the platform maximum before deciding whether to clamp the request. MFC after: 2 weeks
e1000: Allow more time for PCH ULP exit Firmware may take up to one second to unconfigure ULP, and affected Lenovo systems have required nearly two seconds. Allow 2.5 seconds before treating the transition as a PHY failure. This extends DPDK commit 7aa4c34581a5 using the field-tested bound from Linux commit 3cf31b1a9eff. MFC after: 2 weeks
e1000: Check PHY control register reads Do not modify a zero-initialized PHY control value when its preceding read failed. MFC after: 2 weeks
igc: fix RX accounting for multi-descriptor packets The receive path adds the running packet length to rx_bytes for every descriptor. A packet spanning descriptors of length l1, l2, and l3 is therefore counted as 3*l1 + 2*l2 + l3. Add each descriptor length once. Single-descriptor accounting remains unchanged. MFC after: 1 week
igc: make AIM counter sampling coherent Sample free-running counters by delta instead of clearing them from the interrupt filter, which can race their producers. Publish byte and packet counts together at the TX and RX doorbells so each sample is coherent. Aggregate every TX ring assigned to the interrupt vector so unequal RX and TX queue counts are safe. Count RX bytes only after a frame is accepted. MFC after: 1 week
igc: synchronize interrupt moderation state Keep the saved EITR value synchronized with hardware across reinitialization. Correct EITR encoding, decoding, and MSI-X register selection, and reject nonpositive fallback rates. Apply the packet-buffer fallback without permanently disabling AIM. MFC after: 1 week
igc: use packet-size AIM Use the packet-size calculation introduced for igb(4) in a69ed8dfb381 and retained there until the iflib conversion in f2d6ace4a684. It derives interrupt holdoff from average packet size, so RSS queue count does not change its behavior. The calculation follows the pre-iflib igb code. Retain igc's normal and low-latency rate caps, and keep the current setting when an interval has no usable sample. MFC after: 1 week
igc: count TSO wire segments in the AIM counters The transmit path bills one packet of ipi_len bytes per request. For TSO that is the whole unsegmented payload, up to 64 KiB, rather than a packet size that appears on the wire. Count the segments the hardware emits and the header carried by each segment. Non-TSO accounting is unchanged. MFC after: 1 week
No functional change intended. Sponsored by: The FreeBSD Foundation MFC after: 1 week
Previously, multiple frames of xfers are split into many tds. In the refactor process, we forget to consider this. The td builder is already allocate with enough numbers of tds. What we need to do is to fill the normal trbs for muiltiple tds when building trbs. PR: https://bugs.freebsd.org/bugzilla/show_bug.cgi?id=297053 Tested by: phk, oleglelchuk@gmail.com Fixes: https://cgit.freebsd.org/src/commit/?id=e0b235ecd4fa ("xhci: Refactor xhci_generic_setup code") Sponsored by: The FreeBSD Foundation Differential Revision: https://reviews.freebsd.org/D58465
It is better to propagate it to pcm_register(), and later to the device drivers, than to simply ignore it and return ENXIO. Sponsored by: The FreeBSD Foundation MFC after: 1 week
Sponsored by: The FreeBSD Foundation MFC after: 1 week
In align_abort() and tag_check_abort(), if we got a fault while in kernel, do not panic if a fault handler has been provided. We may get such a fault when trying to read or write userland data, it can at least happen with _umtx_op() if an unaligned pointer is provided. Instead, just let the fault handler deal with it. MFC After: 1 week Approved by: andrew Differential Revision: https://reviews.freebsd.org/D58426
specialreg.h is the tree's MSR registry and already carries the Intel RAPL group. Add the AMD RAPL package/core energy and unit MSRs here so the hwpmc RAPL class can reference them without a private driver copy. Use the names Linux's msr-index.h gives these registers. Reviewed by: mhorne, adrian, Ali Mashtizadeh <ali@mashtizadeh.com> MFC after: 3 days Sponsored by: AMD Differential Revision: https://reviews.freebsd.org/D58027
Add hwpmc_rapl.c/.h implementing PMC_CLASS_RAPL, a read-only system-scope class modeled on TSC and wired into x86 AMD and Intel MD init. A per-vendor MSR table covers AMD/Hygon and Intel; energy is reported in microjoules, with the Intel server 2^-16 J DRAM unit handled and 32-bit wraps recovered into a 64-bit accumulator. The overflow guard follows the PMC lifetime: armed on the first allocated PMC, callout_drain()d on the last release, and each tick only rendezvouses CPUs holding one. Per-CPU spin locks guard the accumulator against torn reads on i386. PMC_CAP_DOMWIDE lets pmcstat(8) allocate one counter per NUMA domain instead of per CPU. Reviewed by: mhorne, Ali Mashtizadeh <ali@mashtizadeh.com> Sponsored by: AMD Differential Revision: https://reviews.freebsd.org/D58028
Register PMC_CLASS_RAPL in libpmc: event table, allocator, class-table descriptor, and the event-name/class-listing lookups, all x86-guarded and modeled on the TSC class. Energy events are read-only and unqualified. The class prefix (RAPL-) supplies the friendly spelling, so pmcstat -S rapl-energy-pkg resolves to the canonical ENERGY_PKG event. Add a pmc.rapl.3 manual page documenting the events, counter scope, the microjoule unit and wrap handling, and the NUMA/package domain mapping; link it from pmc.3. Reviewed by: mhorne Discussed with: Ali Mashtizadeh <ali@mashtizadeh.com> Sponsored by: AMD Differential Revision: https://reviews.freebsd.org/D58029
Neither of these options are checked in the file and cdefs.h should not be included explicitly. No functional change. Sponsored by: The FreeBSD Foundation
amd_allocate_pmc() chose the pmu-events code path whenever pmc_cpuid was non-empty, and rejected any allocation lacking PMC_F_EV_PMU. But pmc_cpuid is set for every AMD CPU, while the pmu-events tables only cover Zen and later. On older families (K8, Bobcat, Jaguar/16h, Bulldozer) libpmc finds no pmu-events entry and falls back to the legacy path, which never sets PMC_F_EV_PMU. Reviewed by: mhorne Approved by: mhorne MFC after: 1 week MFC to: stable/14, stable/15 Sponsored by: Netflix Differential Revision: https://reviews.freebsd.org/D58468
{em,igb}_determine_rsstype() mapped only the TCP and bare-IP RSS descriptor
types; the UDP types returned M_HASHTYPE_NONE.
The hardware does hash UDP, but with a NONE hashtype iflib skips its
flowid-based TX queue spread, so all forwarded UDP egressed on a single queue
and serialized transmit on one core.
Add the three UDP cases (IPV4_UDP, IPV6_UDP, IPV6_UDP_EX) so egress spreads
across all TX queues.
Reviewed by: kbowling, gallatin
Approved by: kbowling
MFC after: 1 week
MFC to: stable/14, stable/15
Sponsored by: Netflix
Differential Revision: https://reviews.freebsd.org/D58513
Add a QUIRK_EMPTY_NAMESPACE_CHANGED_LOG quirk which indicates that the nvme controller may not properly populate the namespace-changed log page. If we receive a NVME_LOG_CHANGED_NAMESPACE page for a device with this quirk and the page is empty, probe all of the namespaces rather than none of them. Reviewed by: imp MFC after: 1 week Sponsored by: Amazon Differential Revision: https://reviews.freebsd.org/D58231
This controller exhibits QUIRK_EMPTY_NAMESPACE_CHANGED_LOG behaviour. A bug report has been filed with the vendor. Reviewed by: imp MFC after: 1 week Sponsored by: Amazon Differential Revision: https://reviews.freebsd.org/D58232
In particular, handle authentication errors due to bad MACs when decrypting packets. Since the current dispatch code assumes synchronous OCF sessions by design, explicitly reject any created OCF session that is not synchronous. Software sessions are always synchronous in practice, so this should be a nop. Approved by: so Security: FreeBSD-SA-26:52.if_wg Security: CVE-2026-58085 Reviewed by: markj Sponsored by: Chelsio Communications
Nexus-attached driver that discovers and parses coreboot's LBIO tables from physical memory. Exposes firmware metadata (version, build info, mainboard, serial config, TSC frequency, CBMEM entries) via sysctl hw.coreboot.*, the firmware console ring buffer via /dev/coreboot_console, and structured CBMEM entry access via /dev/cbmem ioctl interface. Tested on: - Qotom Q535G6 (Kabylake) - Intel NUC D54250WYK (Haswell) - Intel NUC D33217GKE (Ivy Bridge) - Dell 3100 2-in-1 (Gabbiter) - Dell 3100 (Fleex) - Lenovo IdeaPad 320s - Lenovo ThinkPad T480 - HP Chromebook 11 G4 - HP Chromebook 11 G5 - HP Chromebook 11 G6 EE - HP Chromebook 14 G4 - HP Chromebook 14 G5 - HP Chromebook x360 11 G1 EE - HP Chromebook x360 11 G2 EE - HP Chromebook x360 14 G1 - Acer C720 - Acer Chromebook 11 - Lenovo N22 Reviewed by: ngie, kib, adrian Differential Revision: https://reviews.freebsd.org/D55649
Intel Apollo Lake SDXC controller reports a Slot Type of "Embedded Slot for One Device" in SDHCI_CAPABILITIES bits, even when the slot is a removable card reader. This caused 48 timeouts before the boot sequence resumed. Reviewed by: imp Differential Revision: https://reviews.freebsd.org/D58467
Update the shared e1000 PF/VF mailbox interfaces for an in-tree igb SR-IOV implementation. Intel FreeBSD igb-2.5.31 and DPDK provide the older PF/VF mailbox baseline. The retained PF mailbox read and explicit unlock operation follow a simple Linux igb parameter addition to make PF mailbox acquisition nonblocking so the driver can retry outside the shared primitive. Treating a CTS-less E1000_PF_CONTROL_MSG as a reset follows DPDK. Sponsored by: BBOX.io
Use each ring's physical queue index for initialization, MSI-X routing, register dumps, sysctls, and debug output instead of assuming that its logical array index is also its hardware index. This is a no-op for the normal queue layout. A later SR-IOV change moves the PF ring to hardware queue num_vfs, so its hardware ID then differs from logical queue zero. Sponsored by: BBOX.io
A non-zero VF device number does not always require ARI. The Intel 82576 and I350 [1] explicitly support a non-ARI layout that places VFs on the next bus. Check every requested VF RID and reject a non-zero device only when it is on the PF bus. This retains the ARI guard for invalid same-bus layouts while permitting the documented second-bus layout. [1] Intel I350 Datasheet, sections 7.8.2.6.1.2, 9.6.4.6 Sponsored by: BBOX.io
pci_iov_config() programs NumVFs before validating the final VF RID layout and allocating all generic resources. A subsequent error ran the driver uninit callback but left the hardware NumVFs register programmed while the software VF count returned to zero. Clear NumVFs in the error path after the driver uninit callback, matching normal SR-IOV teardown ordering. This prevents stale hardware state after a failed configuration and permits a clean retry. MFC after: 1 week Sponsored by: BBOX.io
Register the 82576 and I350 VF PCI IDs under a separate igbv driver while continuing to share the igb datapath implementation. Follow the ixv driver split and give the VF context IFLIB_IS_VF so iflib does not apply the PF SR-IOV detach guard to a child VF. Program VTIVAR_MISC in the VF low byte so mailbox and reset notifications reach the VF admin vector. The split will become increasingly obvious as bug fixes land, trying to bias everything with if (sc->vf_ifp) everywhere is error prone in two directions. This breaks existing naming/configurations and cannot be MFCed as-is. I have no plans of adapting it to prior branches at the moment but it may be possible. Relnotes: yes Sponsored by: BBOX.io
I350 loopback receive descriptors report VLAN tags byte-swapped for both PFs and VFs. The receive path handled the PF device types but omitted e1000_vfadapt_i350, causing an admitted VF VLAN packet to be delivered untagged to the VF parent. Include the I350 VF type in the existing correction. This matches the dedicated IGB_RXQ_FLAG_LB_BSWAP_VLAN handling in DPDK igbvf. MFC after: 1 week Sponsored by: BBOX.io
The register-dump sysctl is installed before iflib allocates the queue arrays and remains visible while they are freed. Return ENXIO outside the queue lifetime instead of dereferencing a NULL or stale array. Sponsored by: BBOX.io
Add the PCI IOV schema and PF control plane for up to seven VFs with one hardware queue per pool. Implement VF mailbox handling, MAC and VLAN assignment, multicast filtering, promiscuity policy, anti-spoofing, malicious-driver recovery, reset replay, and queue lifecycle management. The basic SR-IOV and VMDq PF implementation follows DPDK Intel e1000 code, including PF pool selection, one queue per pool, mailbox dispatch, and VF enablement. Intel FreeBSD igb-2.5.31 supplies the older driver baseline. Linux igb and the Intel SDMs clear up lifecycle, isolation, reset, and family-specific details absent from DPDK. Enabling IOV requires the PF to attach with one TX and RX queue. Systems whose defaults select RSS queues must set the documented iflib queue override tunables before attach. Only 82576 and I350 support SR-IOV in silicon. The series has been extensively tested on I350, including thowing boundaries at the PCI BAR that shipping drivers will never. Still, think carefully before reaching for this in critical environments. Relnotes: yes Sponsored by: BBOX.io
Disable each igb-class transmit and receive queue and flush before changing its descriptor-ring registers. Restore the head and tail indices that Intel documents as surviving a VF reset. Use the igb queue-enable control instead of programming legacy TXDCTL granularity, low-water, and reserved bits that do not belong to the 82575 and later. Sponsored by: BBOX.io MFC after: 1 week
igbv: Isolate VF policy and validate its registers Give igb virtual functions a separate ifdi method table and move VF-specific attach, reset, queue, interrupt, and diagnostic policy to if_igbv.c. Keep shared descriptor-ring mechanisms in if_em.c. Derive VF identity from IFLIB_IS_VF and assert that hardware identification agrees. Under INVARIANTS, validate normal VF CSR accesses against the sparse 82576 and I350 VF register maps. Stop shared setup from accessing PF-only controls. Require MSI-X and defer VF sysctls until attach succeeds so failed attachment cannot leave handlers pointing at freed driver state. Advertise only VF capabilities, run adaptive moderation without the PF receive-buffer guard, enable SRRCTL.DROP_EN, and provide a VF-safe diagnostic register view. The moved implementation is the existing FreeBSD code. Register model was cross-checked against the Intel datasheets and other Intel drivers. Sponsored by: BBOX.io
igbv: Improve VF mailbox and status behavior Treat VF media as fixed 1000baseT full duplex and report PF not ready and generated MAC fallback states during attach. After a successful reset handshake, reconcile a PF rejected MAC back into the ifnet. If the PF is unavailable, defer MAC, multicast, VLAN, LPE, and promiscuity replay until CTS is restored. Track a rejected VLAN removal separately so leaked traffic remains tagged until reset proves that the stale hardware filter is gone. Baseline VF counters at attach, collect the four loopback packet and octet counters with rollover-safe deltas, and account software RX checksum offload results. Preserve accumulated statistics across PF resets by rebasing the raw hardware counters, and sample them while physical link is down because VF loopback can remain active. Retain the 82576 VFMPRC hardware statistic, but do not read it on I350 VFs because specification update errata 31 says it is unavailable. Clear PF owned flow control state and reset a link down VF when queued transmit descriptors must be flushed. Sponsored by: BBOX.io
igbv: Support secondary unicast filters Support the Linux igbvf secondary-MAC mailbox subprotocol, used by Linux guests running MacVTap. Replay up to three non-primary unicast addresses after reset and whenever the address list changes, subject to PF allow-set-mac policy. Sponsored by: BBOX.io
igb: Stop writing the legacy TADV register TADV is an em-class interrupt delay register and is absent from the 82575 and later register model. The igb attach path does not expose or initialize that control, but transmit initialization still wrote its zero valued storage into a reserved queue-window offset. Apply the same igb_mac_min boundary already used for TIDV and the absolute-delay sysctls. MFC after: 1 week Sponsored by: BBOX.io
igb: Update only changed IOV multicast hashes Build the aggregate PF/VF multicast bitmap in software and compare it with the e1000 MTA shadow. Write only registers whose desired value changed, while forcing a complete write after PF reset invalidates the hardware table. This bounds alternating VF multicast updates without NACKing them. Linux igbvf and DPDK ignore multicast reply status, so a command-rate limiter could otherwise acknowledge configuration while leaving hardware state stale. Sponsored by: BBOX.io
igb: Update only changed IOV VLAN filters Keep the full VFTA/VLVF software recomputation and clear-map-set ordering, but compare each phase against the authoritative old value. Write only VFTA words and VLVF slots whose effective contents change. I350 uses its software VFTA shadow because erratum 20 makes live reads unreliable; an invalid shadow forces a complete clear before sparse restoration. 82576 continues to diff against live VFTA reads. Add SDT probes for every logical write phase and the final software images so hardware tests can verify exact elision counts. On my I350 DUT, the old full table path averaged 819 us across 31 VLAN removals versus about 79 us for the PF statistics sweep. Sponsored by: BBOX.io
igb: Rate-limit VF VLAN rebuild requests Give each VF a burst of 64 VLAN additions and refill it at eight additions per second. Removals remain unrestricted, idempotent requests consume nothing, and trusted PF-wide initialization replenishes the burst while guest resets do not. Checks VLVF capacity before charging a token. Do not apply this policy to multicast requests because Linux igbvf and DPDK ignore their reply status; aggregate MTA write elision bounds those updates instead. Sponsored by: BBOX.io
Mailbox and link interrupts share iflib admin service with the periodic timer. Mark timer-driven passes explicitly and run the hardware statistics sweep only for those samples instead of repeating 66 PF MMIO reads for every VF mailbox message. DTrace on the I350 DUT measured the PF sweep at about 79 us on average. The normal hz/2 timer continues to extend clear-on-read counters safely; exported counters may trail hardware by up to 500 ms. Sponsored by: BBOX.io
A PF mailbox NACK does not distinguish the SR-IOV VLAN request rate limit from permanent VLVF exhaustion. Preserve desired VLAN membership and retry four additions per 500 ms timer tick, matching the PF sustained allowance. Bound the whole recovery batch to eight seconds from its first failure and consolidate restore diagnostics, so a full table cannot create a permanent mailbox poller or repeated per-VID log bursts. Sponsored by: BBOX.io
Pass the VF generation through the CSR accessors so the validator can distinguish the sparse 82576 and I350 register maps. Admit the queue-zero RXCTRL, TXCTRL, TDWBAL, TDWBAH, and VFPSRTYPE registers exposed by both families. 82576 exposes VFMPRC at 0xf3c. I350 erratum 31 makes its corrected 0xf38 address inaccessible to a VF, so reject both I350 spellings while retaining read access on 82576. Sponsored by: BBOX.io
82576 and I350 VFLR leave the VF queue configuration unchanged. A VF can program transmit head write-back and leave its DMA destination for a later VF owner; mainstream VF drivers do not overwrite TDWBAL/H. Disable every receive and transmit queue assigned to the VF, wait for the enable bits to clear, then clear SRRCTL, PSRTYPE, RXCTRL, TXCTRL, and TDWBAL/H. Spin briefly for the normal transition, then sleep at 100 microsecond intervals with an approximately 1 ms bound. This prevents a VF that keeps asserting QUEUE_ENABLE from busy-waiting the PF context lock for 10 ms. If a queue does not quiesce, leave the VF disabled and NACK its reset rather than programming an active queue. Rate-limit this diagnostic independently from mailbox and malicious-driver notifications. I350 maps pool n to queue n. 82576 assigns physical queues n and n+8 to VF n, so sanitize both queues while clearing per-pool PSRTYPE once. An incoming VF initializes its active ring base, head, and tail while enabling each queue. This is also required by malicious-driver recovery, which deliberately does not assert VTCTRL.RST because doing so would discard the VF's admin-vector routing before the PF can notify it. This implements Software Clarification 3 from the 82576 and I350 specification updates. Sponsored by: BBOX.io
82576 and I350 VFLR leave queue configuration unchanged. A previous VF owner can therefore leave a transmit head-writeback DMA destination and other queue policy for the next guest. After each reset attempt, disable all exposed VF queues and wait for their enable bits to clear before clearing SRRCTL, VFPSRTYPE, RXCTRL, TXCTRL, and TDWBAL/H. Spin briefly and then sleep until the bounded queue-disable deadline. iflib cannot report initialization failure and marks an interface running after its init callback returns. On sanitation failure, keep interrupts disabled and use the deferred admin task to clear RUNNING. Retry after 100 and 500 ms; after three total failures, leave the interface down until another administrative initialization starts a new bounded attempt set. igbv uses queue zero on both families, but 82576 exposes a second VF queue whose retained state must also be cleared. Extend the INVARIANTS register validator for only those queue-one CSRs and only on 82576. This implements the VF side of Software Clarification 3 from the 82576 and I350 specification updates. It also means an igbv guest does not depend on its PF to sanitize a previous VF owner's state. Sponsored by: BBOX.io
amd64: do not allow to set reserved bits in MXCSR for ptrace(PT_SETFPREGS) Also do not mask bits in the mxcsr_mask. It is ignored by FRSTOR/XRSTOR. Reported by: markj Reviewed by: jhb, markj Sponsored by: The FreeBSD Foundation MFC after: 1 week Differential revision: https://reviews.freebsd.org/D58548
ptrace: Propagate errors from set_fpregs() Otherwise ptrace(PT_SETREGSET) will not return errors to userspace. Fixes: https://cgit.freebsd.org/src/commit/?id=cef05c5a62ba ("amd64: do not allow to set reserved bits in MXCSR for ptrace(PT_SETFPREGS)") Reviewed by: kib Differential Revision: https://reviews.freebsd.org/D58577
Reported by: jhb Reviewed by: jhb, jrtc27 Sponsored by: The FreeBSD Foundation MFC after: 1 week Differential revision: https://reviews.freebsd.org/D58550
Some conventional PCI e1000 configurations hang when given DMA addresses above 4 GB, particularly on systems using AMD HyperTransport-to-PCI bridges. Linux has restricted e1000 to DMA32 in PCI mode since 2011 for the same failure class in commit e508be174ad36b0cf9b324cd04978c2b13c21502. Set iflib's DMA width after determining the negotiated bus type. This covers descriptor and packet-buffer mappings while preserving 64-bit DMA for PCI-X and PCIe devices and providing a conditional tunable. PR: https://bugs.freebsd.org/bugzilla/show_bug.cgi?id=297064 Reported by: Alexander Leidinger <netchild@FreeBSD.org> Tested by: Alexander Leidinger <netchild@FreeBSD.org> MFC after: 1 week
Add register/bit definitions for the L1 PM substates capability (PCIZ_L1PM) to pcireg.h. Signed-off-by: Michael Adler <madler@tapil.com> MFC after: 1 week Pull-Request: https://github.com/freebsd/freebsd-src/pull/2318
I226 parts advertise support for the PCIe L1.2 link substate, but a
hardware erratum makes the exit latency from that low-power state
longer than the packet buffer can absorb under load. This stalls the
inbound packet stream. Disabling ASPM system-wide (BIOS or OS ASPM
policy) does not fix it. The L1.2 enable bit must be cleared directly
in the device's own PCIe L1 PM extended capability.
Add igc_is_device_id_i226() to identify affected parts and
igc_disable_broken_aspm_l1_2() to clear the ASPM L1.2 enable bit
on attach and after resume, since PCIe config space can be
reset across a suspend/resume cycle.
Adapted from the Linux igc driver:
0325143b59c6 igc: disable L1.2 PCI-E link substate to avoid
performance issue
1468c1f97cf3 igc: fix disabling L1.2 PCI-E link substate on I226
on init
Signed-off-by: Michael Adler <madler@tapil.com>
PR: https://bugs.freebsd.org/bugzilla/show_bug.cgi?id=279245
Reviewed by: Jim Thompson
MFC after: 1 week
Pull-Request: https://github.com/freebsd/freebsd-src/pull/2318
ixgbe: isolate VF reset state IXGBE_VF_INDEX() selects a 32-VF register bank. PFMBMEM() selects one mailbox per VF, while ixgbe_toggle_txdctl() calculates queue offsets from a VF number. Passing the bank index aliases VF1-31 to VF0 and VF32-63 to VF1. Resetting one VF can therefore clear the peer mailbox and leave its transmit queues disabled. The VF raises its reset event before posting its mailbox request. The PF checks reset events before mailbox messages. If both are pending, clearing PFMBMEM during generic reset handling can erase the request before ixgbe_read_mbx() consumes it. Clear the mailbox only from the reset-message handler after the request has been read. Use the VF number for queue toggling and document that API contract. MFC after: 1 week
ixgbe: respect peer mailbox ownership A VF currently treats an existing VFU bit as a successful acquisition, while the PF checks its own PFU bit before claiming the mailbox. Check both the local and peer ownership bits before setting local ownership. This prevents same-side callers from sharing the mailbox and avoids an acquisition attempt while the peer owns it. VFLR does not clear VFMAILBOX.VFU. Clear stale VF ownership and cached mailbox status after the reset indication settles and before sending the reset request, so the ownership check cannot strand a reinitialized VF. Adapt only the live ownership checks from Intel ix 3.4.39. Do not import its upgraded-mailbox changes, which are not active in FreeBSD. Obtained from: Intel ix 3.4.39 MFC after: 1 week
ixgbe: fail fast on VF-held PF mailboxes The active PF mailbox operations use the legacy helpers. The mailbox API import changed check_for_msg into a read-only probe and added up to 2,000 500-microsecond lock retries. If a VF leaves VFU set, the PF cannot acquire the lock, busy-waits for up to one second, and leaves VFREQ pending so the delay can repeat. Give the legacy checker its old consume-on-check behavior so a failed read does not leave VFREQ asserted. If VFU is already set, fail immediately instead of retrying, while preserving retries for PF-side contention. Do not force RVFU, which would discard peer transaction state. MFC after: 1 week
ixgbe: enforce configured VF anti-spoofing The SR-IOV schema advertises MAC anti-spoofing and enables it by default, but the VF configuration was never consumed and the hardware policy remained disabled. Record the configured policy and apply MAC and VLAN anti-spoofing throughout VF initialization and reset. On X550-family devices, also protect the LLDP and flow-control Ethertypes and enable per-VF spoof-event accounting. Remove the driver-owned state during SR-IOV teardown. Adapt the anti-spoof configuration lifecycle used by igb(4) in a2ed165f0049 to the ixgbe hardware controls. MFC after: 1 week Relnotes: yes
ixgbe: preserve VLAN ownership with SR-IOV The VF VLAN capability is checked but never granted, and no SR-IOV configuration property exposes the existing default-VLAN support. PF VLAN updates also replace VFTA registers from a PF-only shadow, erasing live VF filters. Expose access VLAN and trunk policy through the IOV schema. Track each VF VLAN as desired state, restore the administrative VLAN after reset, and use the native VLVF helper for incremental PF and VF ownership changes. Keep VLAN filtering enabled while SR-IOV is active. When PF hardware filtering is disabled, admit every VLAN to the PF without bypassing per-pool VF isolation. Reconstruct VLVF and the shared VFTA from PF and VF desired state after reset or a filtering-mode transition, and restore PF-only state on teardown. When the last VF leaves a VLAN still owned by the PF, free its VLVF slot while retaining the shared VFTA bit. This prevents a trunk VF from exhausting the 64-entry VLVF table by cycling VLAN memberships. Adapt the VLAN ownership model introduced for igb(4) in a2ed165f0049 to ixgbe's native VLVF machinery. Match Linux receive semantics by exposing a stripped VLAN tag only when that VID was registered by the VF. A PF-assigned port VLAN is an administrative tag and must be delivered to the VF as untagged traffic; otherwise the stack dispatches it to a nonexistent VLAN interface and access-VLAN receive traffic is blackholed. MFC after: 1 week Relnotes: yes
ixgbe: Preserve priority-tagged traffic with SR-IOV VID 0 carries only 802.1p priority and does not identify VLAN membership. Keep VFTA bit zero in the persistent PF shadow table so reset and SR-IOV replay admit priority-tagged frames while VLAN filtering is enabled. In virtualization mode, also reserve VLVF slot zero and restore PF and eligible VF pool memberships. A VFTA hit alone admits the tag globally but does not deliver it to the correct pools. This matches the priority-tag treatment in em/igb. MFC after: 1 week
ixgbe: enforce VF promiscuity and multicast policy The allow-promisc IOV property is advertised but ignored, and the PF rejects the xcast request used by modern VFs. Negotiate mailbox APIs 1.2 and 1.3, implement pool-scoped xcast modes, and require allow-promisc for requested all-multicast or unicast-promiscuous modes. The VF mailbox can carry only 30 multicast hashes. When ixv has a larger list, request the API 1.2 all-multicast xcast mode instead of extending the legacy SET_MULTICAST message. The PF grants that fallback only to VFs configured with allow-promisc; otherwise ixv reports that only the first 30 addresses are active. Reset xcast state with the VF and have ixv replay the mode implied by its interface flags after multicast updates. Follow DPDK's ixgbe API 1.2/1.3 xcast contract, with allow-promisc policy adapted from igb(4) in a2ed165f0049. MFC after: 1 week Relnotes: yes
ixgbe: implement VF secondary MAC filters The PF advertises the legacy SET_MACVLAN mailbox request but always rejects it. The request installs secondary unicast addresses. Allocate an owned RAR pool for VF secondary addresses, reserve low entries for PF filters, and place VF-primary addresses at the top of the usable RAR range. Reject address collisions and cap each VF at three secondary filters so one guest cannot exhaust the shared table. Clear secondary filters on VF or PF reset and on SR-IOV teardown. This hardware can anti-spoof only the VF primary source address. Reject secondary filters while MAC anti-spoofing is configured, so installing them requires an explicit administrative policy choice. Report optional filter-table allocation failure without disabling SR-IOV. Adapt the owned-RAR allocation and reset-cleanup model from igb(4) in a2ed165f0049 to DPDK's ixgbe SET_MACVLAN mailbox semantics. MFC after: 1 week Relnotes: yes
ixgbe: Preserve VF jumbo frame size across PF resets sc->max_frame_size represents the largest frame requested by the PF or an active VF. The MTU callback replaces it with the PF frame size, so a subsequent reinitialization can program MHADD below an active VF's jumbo-frame request. Recompute the aggregate before hardware initialization and use it when programming MHADD. Recompute after each VF LPE request as well, so a reduced request can lower the hardware limit when no other function needs the previous value. MFC after: 2 weeks
ixgbe: Restore missed packet accounting missed_rx and total_missed_rx are never populated. As a result, the GPRC erratum workaround does not remove missed packets and iqdrops always remains zero. The rx_missed_packets sysctl and input-error total also expose only MPC bank zero. Read and accumulate all eight MPC banks. Use the interval total to correct GPRC and the cumulative total for iqdrops, input errors, and the aggregate sysctl. This matches DPDK's coverage of the hardware banks. MFC after: 2 weeks
ixgbe: Validate EEPROM checksum section bounds The generic checksum walker trusts NVM section pointers and lengths and iterates with a 16-bit index. A corrupt section that crosses the end of the EEPROM can wrap the index and leave the driver in an effectively unbounded read loop during attach. Validate each non-empty section against the discovered EEPROM word size before reading it, and use widened arithmetic for the inclusive end and iterator. MFC after: 2 weeks
ixgbe: Compare flow control against requested mode The flow-control sysctl represents the configured policy, while current_mode is the mode negotiated with the link partner. Comparing a new request with current_mode can needlessly reprogram an unchanged policy or skip a requested policy change that happens to match the current negotiation result. Compare with requested_mode before deciding that no update is needed. MFC after: 2 weeks
if_foreach_llmaddr() adds each callback return value to its running count. Returning the incremented count made the address indices grow as 0, 1, 3, 7, and so on, eventually writing beyond the multicast address array. Return one address per callback and stop copying when the array is full, matching the ixv-1.6.12 driver. Fixes: https://cgit.freebsd.org/src/commit/?id=ff06a8dbb677 ("Mechanically convert ixgbe(4) to IfAPI") MFC after: 1 week
DPDK commit message net/ixgbe/base: add missing buffer copy for ACI Add the missing buffer copy in ixgbe_aci_send_cmd(). The retry path saves the original descriptor and allocates storage for the command buffer so both can be restored before another attempt. It did not copy the original command buffer into that storage. Fixes: https://cgit.freebsd.org/src/commit/?id=25b48e569f2f Cc: stable@dpdk.org Signed-off-by: Dan Nowlin <dan.nowlin@intel.com> Signed-off-by: Yuan Wang <yuanx.wang@intel.com> Acked-by: Bruce Richardson <bruce.richardson@intel.com> Obtained from: DPDK (37239792b0) MFC after: 1 week
DPDK commit message net/ixgbe: fix flow control frame byte adjustment LXONTXC and LXOFFTXC are 32-bit counters for transmitted XON and XOFF packets. Their deltas are summed and used to adjust the transmitted packet and byte counters. Perform the addition in 64 bits so it cannot wrap before the result is used for the byte adjustment. Found by Linux Verification Center (linuxtesting.org) with SVACE. Fixes: https://cgit.freebsd.org/src/commit/?id=af75078fece3 ("first public release") Cc: stable@dpdk.org Signed-off-by: Daniil Iskhakov <dish@amicon.ru> Acked-by: Bruce Richardson <bruce.richardson@intel.com> Obtained from: DPDK (bdf8608559) MFC after: 1 week
DPDK commit message net/ixgbe/base: fix unchecked return value Check the return value from ixgbe_read_eeprom() before using the control word to configure link disable during D3. Fixes: https://cgit.freebsd.org/src/commit/?id=b7ad3713b958 ("ixgbe/base: allow to disable link on D3") Cc: stable@dpdk.org Signed-off-by: Barbara Skobiej <barbara.skobiej@intel.com> Signed-off-by: Anatoly Burakov <anatoly.burakov@intel.com> Acked-by: Bruce Richardson <bruce.richardson@intel.com> Obtained from: DPDK (eb3684b191) MFC after: 1 week
FreeBSD's I2C helper already retries failed transactions. Limit this new outer loop to successful reads with an invalid identifier so that retry budget is not multiplied. DPDK commit message net/ixgbe: retry misbehaving SFP read Some XGS-PON SFPs ACK I2C reads and return uninitialized data while their microcontroller boots. A bogus identifier can cause an otherwise working module to be marked unsupported. Retry the identifier read several times, checking for both successful I2C completion and a valid SFP identifier. Signed-off-by: Stephen Douthit <stephend@silicom-usa.com> Signed-off-by: Jeff Daly <jeffd@silicom-usa.com> Reviewed-by: Haiyue Wang <haiyue.wang@intel.com> Obtained from: DPDK (774263bb4e) MFC after: 1 week
ixgbe: fix unaligned access in ixgbe_update_flash_X550() ixgbe_host_interface_command() treats its buffer as a u32 array. The local union contained only byte-sized fields, giving it one-byte stack alignment and allowing unaligned accesses on strict-align systems. Add a u32 member to the union to provide the required alignment and pass that member to ixgbe_host_interface_command(). No functional change is expected on x86. Obtained from: Intel ix 3.4.39 MFC after: 1 week
ixgbe: avoid signed overflow in pause time calculation pause_time is promoted to signed int before multiplication. Its default value of 65535 multiplied by 65537 exceeds INT_MAX and triggers UBSAN, even though the result is assigned to a u32. Make the multiplier unsigned so the calculation has the intended u32 semantics. Linux commit 3b70683fc4d6 reported the failure in the generic path and used the same mechanical correction. The 82598-specific flow control operation contains the identical expression, so correct it as well. MFC after: 1 week
ixgbe: reject VF requests before CTS A VF that sends a non-reset request before completing reset negotiation has not received CTS. The PF ignores the request but currently reports success, leaving the VF with a false view of the programmed state. Return failure for the ignored request. This restores the behavior lost when the mailbox helpers were renamed. Fixes: https://cgit.freebsd.org/src/commit/?id=36c516b31136 ("ixgbe: update if_sriov to use the new mailbox apis") MFC after: 1 week
ixgbe: check negotiated API for VF queue query The GET_QUEUES handler switches on msg[0], which contains the mailbox command rather than the negotiated API version. It therefore cannot reject API 1.0 or an unnegotiated VF as intended. Switch on the API version stored for the VF. MFC after: 1 week
ixgbe: complete PF cleanup after VF FLR The 82599, X540, and X550 documentation identifies VF registers which retain state across VFLR and must be reconfigured before a VF is reused. The VF reset path already initializes its queue-owned registers, but the PF only cleared VF mailbox memory and transmit head write-back addresses after a cooperative mailbox reset. A bare hardware VFLR therefore left both behind on affected devices. Move TDWBA cleanup into the common reset path. Clear CTS when VFLR invalidates the mailbox session, and accept only VF_RESET during the reset pass before restoring VF traffic. Clear VFMBMEM through the PFU/VFU semaphore. Recheck VFREQ while holding PFU so a reset event cannot erase a request posted between the initial mailbox check and the clear. Dispatch an already-read message even if the residual clear fails, but keep cleanup pending until a synchronized clear succeeds. Retry cleanup in the same admin pass after a failed message read or clear. The 82599 also retains VFMAILBOX.VFU across VFLR. Leave a VF-owned mailbox intact initially so a live post-reset writer can finish. Retry cleanup from the admin timer and, after a two-second grace period, use PFMAILBOX.RVFU only when VFU remains set and no request has been posted. Clear the mailbox under PFU afterward. This recovers an abandoned pre-reset owner without sleeping under the iflib context lock or immediately stealing from a new reset request. Suppress mailbox dispatch once iflib has cleared IFF_DRV_RUNNING so a pending reset request cannot re-enable VF traffic inside the PF stop path. Periodically sample aggregate VFREQ, VFACK, and VFLR registers, masked to active VFs, so work suppressed across a stop/restart and a bare 82599 VFLR without EICR_MAILBOX are both discovered without another interrupt edge. MFC after: 2 weeks
ixgbe: recover from X550 malicious-driver events The shared X550 code provides malicious-driver detection, event decoding, and per-pool recovery operations, but the PF never enables or services them. A malformed VF descriptor can therefore go undetected and avoid the per-pool recovery path supplied by the MAC. Configure IOV state while VF DMA remains disabled, then enable MDD and activate the VFs only after PF queue initialization is complete. On an MDD event, withdraw mailbox CTS and gate the VF pool through PFVFTE and PFVFRE. Retain the per-queue WQBR blocks until the VF enters a new reset epoch; PFVFTE can still permit descriptor fetches into the internal queue, so releasing WQBR early would allow a hostile VF to retrigger MDD before it resets. Send the non-CTS reset notification after servicing the VF mailbox. Let a posted VF request win mailbox arbitration, defer notification if the pass produced a response, and retry failed notifications from the periodic admin pass. Poll WQBR so recovery does not depend on another mailbox interrupt edge, while suppressing already-fenced pools. Latch a PF reset request until the next hardware initialization. The X550 datasheet defines every bit of WQBR_RX and WQBR_TX as a queue bit, so an all-ones value is valid. Reject it only when IXGBE_STATUS, which has reserved-zero bits, also reads as all ones and confirms dead MMIO. Temporarily disable MDD around live multiqueue SRRCTL drop-mode updates, which hardware otherwise reports as queue-context changes. Serialize that window with the iflib context lock and resample pending work after MDD is restored. Apply the per-pool recovery model used by igb(4) in a2ed165f0049 to the existing DPDK-derived X550 hooks. The same register interface is documented for X552 and X553, so cover the entire X550 family. Document that VF traffic remains disabled until the reset handshake completes. MFC after: 2 weeks Relnotes: yes
ixgbe: force receive drops on every VF queue PFQDE is indexed by absolute receive queue, but the driver programs one index per VF. Only the first quarter or half of the VF queues therefore have queue-drop isolation, depending on the virtualization mode. The flow-control path can also clear those bits even though SR-IOV requires them independently of the PF pause policy. Program every queue in a VF pool before enabling receive for that VF. For an X550-family VF with an administrative port VLAN, also hide the VLAN tag as the hardware requires. Keep PF flow-control changes confined to the PF SRRCTL registers, and clear the VF queue settings when SR-IOV is torn down and the queues can be reassigned to the PF. MFC after: 2 weeks
ixgbe: Avoid a signed shift while assembling the PBA number The EEPROM word is promoted to signed int before the left shift when the cast is applied to the complete expression. Cast the word first so all 16-bit values are shifted as unsigned data. This is the ixgbe counterpart of the e1000 correction imported from DPDK commit b932270c66. MFC after: 2 weeks
ixgbe: Avoid signed overflow in LED register masks LED index three shifts the blink bit into bit 31. Convert the base to the register width before shifting so the operation is unsigned. This is the ixgbe counterpart of the e1000 correction imported from DPDK commit 214cb0d7f1. MFC after: 2 weeks
ixgbe: Use unsigned register bitmap shifts VLAN, VMDq, and VF reset bit indices can reach 31. Use unsigned values when constructing their 32-bit register masks so the shifts do not operate on signed integers. MFC after: 2 weeks
Clear ROMPE for an empty list and enable it only for a nonempty list. FreeBSD already clears ROMPE when resetting a VF, so that part of the DPDK change is not needed. DPDK commit message net/ixgbe: fix over using multicast table for VF VMOLR.ROMPE allows a VF to receive packets matching the shared multicast table. Leaving it enabled after the VF removes its last multicast address lets PF or peer-VF table entries continue selecting that VF. Signed-off-by: Wei Zhao <wei.zhao1@intel.com> Acked-by: Qi Zhang <qi.z.zhang@intel.com> Obtained from: DPDK (dc5a6e7422) MFC after: 1 week
ixgbe: fix host interface timeout detection The host-interface polling loop was scaled from milliseconds to microseconds, but its terminal test was left using the unscaled timeout. Completion at that intermediate iteration can be reported as a timeout, while actual expiry is not recognized and can accept stale status. Test against the scaled loop bound used by the polling loop. Fixes: https://cgit.freebsd.org/src/commit/?id=f46d75c90f5f ("ixgbe: improve MDIO performance by reducing semaphore/IPC delays") MFC after: 1 week
ixgbe: avoid signed shift when assembling ETrack ID Obtained from: Intel ix 3.4.39 MFC after: 1 week
ixgbe: dispatch PBA string reads through EEPROM ops E610 installs a device-specific PBA string reader, but the public API always calls the generic implementation. Dispatch through the EEPROM operation table so device overrides are honored. Initialize the generic operation for devices that use the ordinary EEPROM representation. Obtained from: Intel ix 3.4.39 MFC after: 1 week
ixgbe: clear VF head write-back state on reset VF reset and FLR do not clear the transmit head write-back address registers. A previous VF driver can therefore leave DMA write-back enabled with a stale address for the next driver instance. After consuming the reset request and disabling the VF queues, clear the address registers for each queue belonging to that VF. Derive the queue count from the active IOV mode so peer queue state is not touched. Linux commit dbf231af81a7 documents the hardware behavior. The FreeBSD implementation follows the local queue mapping and register interfaces. MFC after: 1 week
ixgbe: Recover legacy VFs from invalid DMA targets 82599 and X540 lack the X550 malicious-driver detector. Detect a VF whose PCI status reports a received master abort while its transmit ring has outstanding descriptors and makes no progress across consecutive samples. Consume the accepted PCI status latch, gate that VF I/O, and recover one pending VF per task pass with round-robin selection. This prevents an unreadable function from starving detection or recovery of other VFs. Save the complete writable VF PCI configuration before FLR, restore it afterward, and verify the hardware-backed Command state. Preserve the first good snapshot and pending state across reset events until restore and verification succeed. Introduce a common I/O-disabled policy bitmask so later quarantine policy can extend traffic gating without duplicating fault-state checks. MFC after: 2 weeks
ixgbe: quarantine repeatedly faulting legacy VFs A guest can reinitialize after a VF function-level reset and repeatedly strand an 82599 or X540 PF with invalid descriptor DMA targets. Count only distinct Received Master Abort events accepted by the qualified transmit-stall detector and quarantine the VF after five events. Preserve quarantine across PF reinitialization, reject reset mailbox requests, and keep transmit, receive, and clear-to-send disabled. Recreating SR-IOV clears quarantine. Expose the affected pools through a read-only bitmap. After a successful quarantine FLR, leave the function in post-FLR configuration, explicitly keep decode and bus mastering disabled, verify the Command register, and refresh its PCI-layer cache so a later restore cannot re-enable the function. This addresses CVE-2021-33061 on 82599. Apply the same bounded-failure policy to X540 as defense in depth; the CVE does not list X540. Intel documents the 82599 issue in: http://iommu.com/datasheets/ethernet/controllers-nics/intel/ixgbe/Intel_82599_Application_Note_655276.pdf MFC after: 2 weeks Security: CVE-2021-33061
ixgbe: Re-enable the SFP laser during initialization ixgbe_if_stop() disables the transmit laser on every 82599 SFP fiber port, but the iflib initialization path did not re-enable it. Re-enable the laser before deferred SFP module setup so interface reinitialization cannot leave either single-speed or multispeed optics dark. The hardware wrapper is a no-op when laser control is unavailable. The placement follows Intel ix-3.4.39; this version deliberately applies to every SFP port affected by the stop path. MFC after: 1 week
ixgbe: Defer ECC recovery to iflib The link interrupt filter performed a full hardware reset in interrupt context. This bypassed iflib stop and initialization, including queue quiescence and restoration of temporary LED state. Record the ECC event in the administrative request mask and ask iflib to perform the reset from its taskqueue. Keep the ECC cause masked until reset so the intermediate admin pass cannot re-enable a sticky condition. Handle ECC independently of Flow Director and in legacy interrupt mode. Remove the redundant EICR write; the filter has already cleared the reported causes. Also remove the accompanying complement-mask update of mac.flags. It set every flag except DOUBLE_RESET_REQUIRED and had no place in ECC recovery. MFC after: 2 weeks
ixgbe: Defer firmware recovery transitions to iflib The firmware-mode callout invoked ixgbe_if_stop() directly. This performed a full device reset without the iflib context lock or the iflib queue lifecycle. It could also poll the E610 firmware command interface from callout context while identification was active. Request an iflib reset from the callout instead. Reject initialization while firmware recovery remains active. This leaves the interface stopped and lets iflib publish that state. Request initialization when firmware exits recovery so an administratively-up interface can recover without operator intervention. MFC after: 2 weeks
ixgbe: Defer E610 thermal shutdown to iflib The E610 firmware event handler invoked ixgbe_if_stop() directly from IFDI_UPDATE_ADMIN_STATUS(). This reset the device without the iflib queue lifecycle and left the interface marked running after its hardware was stopped. Request an iflib reset instead. Fail the automatic initialization once so the reset transaction stops the interface and publishes that state. A later operator-requested initialization remains possible, matching the previous recovery policy without bypassing iflib. MFC after: 2 weeks
Fixes: https://cgit.freebsd.org/src/commit/?id=fdc1f3450634 ("x86: change signatures of ipi_{bitmap,swi}_handler() to take pointer") MFC after: 1 week Sponsored by: The FreeBSD Foundation
acpi_probe_child() keeps PCI link devices, the RTC, and docking stations enabled even when _STA reports them not present, but skipped acpi_parse_resources() for them. With an empty resource list, resource-based hint matching (BUS_HINT_DEVICE_UNIT) cannot wire such a device to its hinted unit, and the hinted ISA device is then created as a duplicate. Modern AMI firmware reports the PNP0B00 RTC as not present while handing timekeeping to the ACPI Time-and-Alarm device. Reviewed by: adrian, jhb Differential Revision: https://reviews.freebsd.org/D58047
Classify I226_LMVP and I226_BLANK_NVM as I226 silicon so they receive the I226-specific ASPM L1.2 workaround. PR: https://bugs.freebsd.org/bugzilla/show_bug.cgi?id=279245 MFC after: 1 week Pull-Request: https://github.com/freebsd/freebsd-src/pull/2318
When allow-set-mac is disabled, the MAC filter validation condition rejects the assigned VF unicast address while allowing any different unicast address. The equality test was accidentally inverted when this code moved to the boolean address helper. Accept multicast and the assigned unicast address, and reject other unicast addresses as intended. Fixes: https://cgit.freebsd.org/src/commit/?id=7d4dceec1030 ("ixl(4): Fix VLAN HW filtering") MFC after: 3 days
The conventional VLAN filter update skipped zero shadow words. Removing the final VLAN represented by a VFTA word therefore left the hardware bit programmed even though the software shadow was clear. Pass the changed word to em_if_vlan_filter_write() and write it even when its new value is zero. Retained nonzero words continue to be replayed as before.
I225 devices can incorrectly enter L1 substates while CLKREQ# is asserted, both while idle and in D3. Disable ASPM and PCI-PM L1.2 on I225 to prevent the resulting packet loss. Keep the I226 workaround ASPM-only because it addresses a separate traffic exit latency observation. PR: https://bugs.freebsd.org/bugzilla/show_bug.cgi?id=265714 MFC after: 4 days
igb: preserve coalesced 82576 MDD events WVBR is read-clear, so reading it from the deferred admin pass loses earlier queue bits when multiple VF malicious-driver events arrive before that pass. Snapshot WVBR in the interrupt filter, translate its staggered queue bitmap to pool bits, and OR observations into software latches for deferred notification and recovery. Retain the one-queue VMDq policy used for mixed-driver safety (the vswitch cannot handle a 2Q guest loopback to a 1Q guest per errata).
igb: drain stale MDD state before interrupt arm IOV policy setup can leave MDDET and its read-clear diagnostic registers populated while the admin vector is masked. Carrying that state across the unmask can suppress the next spoof-event edge. Mark initialization for a one-shot drain and consume LVMMC, WVBR when applicable, and ICR immediately before EIMS/IMS arms the vector. Preserve the synthetic link-status cause across the arm-time ICR read, and clear the one-shot latch at reset preparation.
igb: recover retained i350 admin interrupts I350 can retain EICR.OTHER with MDDET and LVMMC asserted while the admin vector and legacy cause remain enabled. The anti-spoof filter continues dropping packets, but no MSI-X is delivered and the spoof diagnostic is lost. Preserve the one-shot setup drain across iflib reset preparation, clear ICR before LVMMC during i350 setup, and kick the enabled admin vector from each admin pass. The synthetic no-cause interrupt stays in the filter and also releases a retained MDDET cause. Keep 82576 drain ordering and stop-time cleanup unchanged.
I225 v1 cannot receive the minimum inter-packet gap required at 2.5 Gb/s. For affected back-to-back links, Intel recommends using a 15-byte transmit IPG instead of 12 bytes. Program TIPG.IPGT to 0xb for pre-v2 I225 devices at 2.5 Gb/s and restore the default at lower speeds. Avoid penalizing fixed I225 and I226 parts. MFC after: 2 weeks
Track RERC separately instead of adding receive errors to the collision count, and read the previously omitted RXERRC register. Include RFC in input errors because CRCERRS does not count bad-CRC runts, implementing the I225 length-error accounting workaround alongside RUC and ROC. Stop treating host transmit MAC discards as receive errors. Expose both RERC and HTDPMC as dedicated MAC statistics so their overlapping counts remain available without corrupting aggregate interface counters. MFC after: 2 weeks
Both these functions have non-static linkage for good reasons, however, their naming may confuse folk when working with crypto(9) code at global scope.
Borrow the e1000 VLAN filter table Ambiguous presence of the feature by Intel was settled by DPDK and emperical testing. MFC after: 2 weeks Relnotes: yes
X550-family malicious-driver detection validates the transmit context selected by a data descriptor with Check Context set. ixgbe sets that bit on every transmit data descriptor, but ordinary PF packets without a VLAN or checksum offload do not create a context descriptor. The empty context then reports an invalid MAC-header length and blocks the PF queue as soon as MDD is enabled. Create the existing context descriptor for every PF packet while SR-IOV is active. This supplies the required MAC-header length and keeps MDD from mistaking normal PF traffic for a malicious-driver event. MFC after: 1 week
Request the reset through iflib and let the admin task perform the stop/init under the context lock, matching what the VF and SR-IOV paths already do. The assertion is compiled out without INVARIANTS, where the same write instead resets the MAC and takes the ICH software flag while the queues stay live and an ioctl or the admin task may be running. While here also remove unnecessary em_if_init uses: iflib_if_init_locked() already runs after IFDI_RESUME and IFDI_MEDIA_CHANGE, so the trailing *_if_init() only added an unstopped IFDI_INIT that the following iflib_stop() undoes. MFC after: 1 week Differential Revision: https://reviews.freebsd.org/D58628
igc_sysctl_eee() and igc_sysctl_dmac() called igc_if_init() directly. Request the reset through iflib instead, and skipping while the interface is down; the new value is picked up by the next init. Unlike e1000, igc has no ASSERT_CTX_LOCK_HELD and no acquire_swflag path, so the defect is silent here rather than an assertion failure. While here also remove unnecessary igc_if_init uses: iflib_if_init_locked() already runs after IFDI_RESUME and IFDI_MEDIA_CHANGE, so the trailing *_if_init() only added an unstopped IFDI_INIT that the following iflib_stop() undoes. MFC after: 1 week Differential Revision: https://reviews.freebsd.org/D58629
gpio: add generic Intel GPIO pin controller framework Add a platform-independent driver framework for Intel GPIO pin controllers found on modern Intel SoCs. The driver accesses GPIO pad registers through ACPI-provided memory-mapped resources and implements the gpio interface [1] including pin enumeration, capability reporting, configuration, and read/write/toggle operations. A common data model of communities and pad groups allows individual SoC-specific drivers to supply their own pad tables and ACPI hardware IDs while sharing all register-level logic. [1] https://wiki.freebsd.org/GPIO Reviewed by: vexeduxr MFC after: 1 week Sponsored by: Beckhoff Automation GmbH & Co. KG Pull Request: https://github.com/freebsd/freebsd-src/pull/2205
gpio: add Intel Alder Lake-N GPIO driver Add a GPIO driver for the Intel Alder Lake-N platform based on the generic intelgpio framework. The driver provides pad group definitions for four GPIO communities covering groups GPP_A through GPP_T, vGPIO and HVCMOS, and matches ACPI hardware IDs INTC1056, INTC1057 and INTC1085. The kernel module build infrastructure and the wiring into files.x86 are included. Reviewed by: vexeduxr MFC after: 1 week Sponsored by: Beckhoff Automation GmbH & Co. KG Pull Request: https://github.com/freebsd/freebsd-src/pull/2205
gpio: add Intel Tiger Lake-H GPIO driver Add a GPIO driver for the Intel Tiger Lake-H platform based on the generic intelgpio framework. The driver defines five GPIO communities with pad groups GPP_A through GPP_K, vGPIO and JTAG, and matches ACPI hardware ID INT34C6. The kernel module build infrastructure and the wiring into files.x86 are included. Reviewed by: vexeduxr MFC after: 1 week Sponsored by: Beckhoff Automation GmbH & Co. KG Pull Request: https://github.com/freebsd/freebsd-src/pull/2205
The VF admin vector carries both link and PF mailbox causes, but the filter schedules the admin task only for link-status changes. Defer administration for every interrupt so reset and control notifications are serviced promptly. MFC after: 1 week
The shared VF set-RAR helper restores hw.mac.addr when the PF rejects a requested address, but ixv ignores the error and leaves the interface link-layer address unchanged. Subsequent initialization repeats the rejected request while the interface appears to use an address the PF will not deliver. Refresh the permanent address returned by the PF after every successful reset handshake. Copy the resulting PF-approved address back to the interface and emit the normal link-layer address notification without re-entering the driver initialization path. This also recovers from a prior mailbox transport failure or a PF-side reassignment. Adapt the igb VF address reconciliation added in a6bb3850e7c6. MFC after: 1 week
The MTA is shared by the PF and all VFs. The VF mailbox handler only ORs new bits, so hashes survive list removal and VF reset. Conversely, PF multicast updates replace the whole table with PF-only state and discard live VF filters. Rebuild the table from the PF list and every active VF whenever either changes. Clear VF multicast state during reset and PF reinitialization, and remove all VF hashes on SR-IOV teardown. Keep the software shadow and multicast control state synchronized, and avoid writes to unchanged MTA registers. Adapt the aggregate desired-state rebuild introduced for igb(4) in a2ed165f0049 and its write-elision scheme from 350211ab1782 to ixgbe's shared MTA. MFC after: 1 week
In preparation of increasing the KVA on powerpc64 to 2TB to mirror amd64's, rework the 64-bit Book-E pmap to not allocate all page table pages at boot time, since that would be a waste of a lot of memory. Instead, allocate all page table pages for the higher levels, leaving the leaves (page directories) for dynamic allocation. This cuts the boot-time page table size down from ~64MB to ~8MB with the current 32GB KVA size, and bumping to 2TB KVA the boot-time page table is still ~8MB instead of ballooning to ~4GB of mostly wasted space.
This reflects what amd64 has, and is needed for using GPUs with large VRAM.
VLAN registration callbacks only update the software shadow, leaving the PF unaware until a later full initialization. Initialization then retries each failed request in a tight loop, while skipping replay entirely when local hardware filtering is disabled. Send additions and removals as soon as the desired state changes, independent of the VF local-filter capability. Replay the desired memberships after reset and retry a bounded batch per timer tick. Stop after the first failure so a silent PF can consume only one mailbox timeout per pass, while a responsive PF can drain several requests. Treat the retry window as a no-progress deadline: advance it when pending work succeeds so a large backlog can drain, but leave entries dormant after a sustained failure. A successful mailbox request wakes a dormant backlog. Dispatch timer-driven retries only while iflib marks the VF running, so a stale timer tick cannot restore PF VLAN state after the stop path resets the VF. Because the callbacks now update the PF or retain failed work for retry, do not restart the VF for VLAN configuration changes. This avoids resetting and flapping the interface for every VLAN addition or removal. Also keep receive VLAN stripping synchronized in both the enabled and disabled cases. Adapt the bounded VLAN reconciliation scheme from igb VF commit fdce3830d9a6 to the ixgbe VF mailbox. MFC after: 1 week
pci: Ignore SR-IOV VFs when tuning MPS The VF Device Control MPS and MRRS fields are reserved and preserved. VF transactions use the PF MPS, so a hardwired VF value must not be used to retune the shared PCIe hierarchy. Document the previously undocumented tuning knob and clarify why a VF may continue to display its reserved hardwired value. This fixes an instant crash/reboot on my Zen3 system with 82599 VFs. MFC after: 1 week
pci: Preserve adjusted PCIe control state The PCI bus changes live capability registers after the initial configuration snapshot has been saved. A later driver reprobe restores that snapshot and can silently undo the adjustment. Update the cached Device Control and Root Control bits together with pcie_adjust_config() writes. Route the persistent Maximum Read Request setter and the bus-owned AER control changes through that helper as well, so they share the same restore semantics as MPS reconciliation. Document the persistent-write contract. Merge only explicitly adjusted bits into the saved image so unrelated or transient bits observed during the hardware read-modify-write cannot become persistent. MFC after: 2 weeks
pci: Reconcile MPS before attaching PCIe devices Reconcile each newly enumerated link as a unit before child drivers attach. Firmware may leave Bus Master Enable set after handoff, so use the bus attachment state rather than that bit to identify the cold phase. Preserve an established hierarchy during rescan and hot-add. Refuse a reduction below a switch because recursive enumeration may already have made a sibling subtree live; lowering only the local port or Root Port would produce an inconsistent path. Report capability and active-use conflicts distinctly. Handle OFW PCI buses that clone the generic enumeration path. MFC after: 2 weeks
pci: Add a hierarchy-wide MPS limit Add a boot-time ceiling for MPS reconciliation. Apply it only while an entire cold-enumerated link can be configured consistently, and leave an established active path unchanged. MFC after: 2 weeks
pci: Optionally disable endpoints with unsafe MPS Keep warn-only behavior as the default. Add an opt-in policy that clears endpoint decoding and bus mastering when a newly discovered function cannot match its active path, while never disabling bridge functions and their subtrees. MFC after: 2 weeks
pci: Permit function-level reset of 82599 VFs Intel 82599 supports FLR on VFs but reports FLR support only in the PF Device Capabilities register. The VF register therefore leaves the FLR Capable bit clear, and pcie_flr() rejects the reset. Intel documents the zeroed VF PCIe capability structure as erratum 35 in the 82599 Specification Update (B0=Yes; NoFix). Add a positive FLR quirk for the 82599 VF. Keep the capability check for every other function, so an unknown nonconforming VF cannot make pcie_flr() report success when its reset request was ignored. SR-IOV requires VFs to support FLR, but a clear capability bit cannot distinguish the 82599's misadvertisement from a VF that fails to implement it. MFC after: 1 week
The IOV callback changes the PF pool, virtualization mode, and hardware queue indices while iflib still considers the old queue layout live. Teardown likewise leaves the software pool and mode at their SR-IOV values. Use iflib stop/mutate/restart transactions for both transitions. Disable VF DMA and PCI VF Enable before queue reuse, let outstanding transactions drain, and restore the non-IOV pool and queue indices on teardown. Remove the redundant driver-local pci_iov_detach() wrapper; iflib already performs that check centrally before the driver detach callback. It may be possible to avoid some restart in the future on this hardware pausing DMA and remapping rings but not pursued yet. MFC after: 2 weeks
A deterministic IOV configuration error currently reaches the driver only after iflib has stopped the PF. The required cleanup restart then causes an avoidable carrier flap. Follow the igb pattern and validate the request in the PCI IOV method before entering the restart transaction. Reject queue layouts wider than the selected virtualization pool before they can alias unrelated 82599 registers. MFC after: 2 weeks
Currently, the FreeBSD driver configures and destroys queues sequentially by issuing individual Admin Queue (AQ) commands. During queue teardown (e.g., interface reset), disabling queues one by one leaves the device in a partially configured state. Because the device does not yet know that the driver is in the process of fully unconfiguring all queues, this intermediate state can trigger transient error logs (such as when queue 0 is disabled while other queues are still active). Modify the driver to use Admin Queue batching for both the creation and destruction of TX and RX queues. Commands are now queued and kicked together, ensuring the queue configuration changes are applied atomically and preventing transient errors from being logged. Signed-off-by: Sujithra Periasamy <sujithra@google.com> Reviewed by: markj MFC after: 1 week Sponsored by: Google Differential Revision: https://reviews.freebsd.org/D58696
A VF's pci_devinfo references its PF's pcicfg_iov for resource bookkeeping, but only the PF implements the SR-IOV capability. pci_cfg_save() and pci_cfg_restore() treated any non-NULL cfg.iov as an owned capability and accessed the PF capability offset in VF configuration space. Saving a VF could therefore replace the shared PF settings with unrelated VF register values. Skip SR-IOV capability save and restore for PCICFG_VF children. The generic PCI and PCIe state of the VF remains preserved. This is also required by drivers that save VF state around a PF-driven function-level reset. MFC after: 2 weeks
When link polling loses mailbox clear-to-send or times out, request an iflib reset instead of continuing with stale VF state. The driver callback runs after iflib samples reset requests, so requeue the admin task to make iflib consume the request on its next pass rather than waiting for an unrelated timer or interrupt. MFC after: 2 weeks
ixgbe: Apply the 82599 D3 link workaround only for D3 ixgbe_stop_mac_link_on_d3_82599() implements the workaround for 82599 erratum 33. It forces incompatible auto-negotiation settings before the device enters D3, and reset clears them when returning to D0. ixgbe_if_stop() is also used for ordinary interface reconfiguration and recovery. Those paths do not enter D3 and should not program this power-management workaround. They continue to stop the adapter and disable the transmit laser. Move the call to ixgbe_setup_low_power_mode(), after ixgbe_if_stop(). This preserves the required ordering for detach, shutdown, and suspend while avoiding the D3 settings during ordinary restarts. MFC after: 2 weeks
ixgbe: Use PF MTU for 82599 VF jumbo policy The shared maximum frame size is raised by VF LPE requests, so it cannot describe the PF MTU when enforcing the 82599 PF/VF jumbo restriction. Consult the PF ifnet MTU instead. Also correct the API 1.1 and later comparison so a jumbo VF is enabled when, and only when, the PF itself uses a jumbo MTU. This matches the policy implemented by DPDK. MFC after: 2 weeks
ixgbe: Quiesce VFs across PF reset Stop VF transmit and receive in hardware, clear PF-side mailbox CTS, and notify active VFs before resetting a PF. A PF reset invalidates VF queue state, so the no-CTS control message makes cooperative VFs discard stale state and renegotiate after the PF returns. The hardware queue gates synchronously prevent further VF DMA. Do not hold the exclusive iflib context lock for a fixed VF-watchdog interval after the reset. Report the PF link transition directly instead of dispatching mailbox work from the stop path, which could otherwise re-enable VF I/O mid-reset. The CTS, PF-control, and VF queue controls follow the reset mechanisms used by DPDK. MFC after: 2 weeks
netmap: Fix driver name handling if_initname() requires the caller to ensure that the lifetime of the interface's name buffer contains that of the ifnet itself. netmap_vi_create() wasn't respecting that; we were instead passing the stack-allocated buffer provided by the ioctl handler. While here, add a check to avoid assuming that the caller-provided buffer is nul-terminated. Reported by: syzkaller Reviewed by: vmaffione MFC after: 2 weeks Sponsored by: The FreeBSD Foundation Differential Revision: https://reviews.freebsd.org/D58676
netmap: Fix a race in kqueue registration We need to acquire the netmap global lock earlier, to avoid racing with the NETMAP_REQ_REGISTER ioctl handler. Reported by: syzkaller Reviewed by: vmaffione MFC after: 2 weeks Sponsored by: The FreeBSD Foundation Differential Revision: https://reviews.freebsd.org/D58677
netmap: Handle overflow when computing ring sizes PR: https://bugs.freebsd.org/bugzilla/show_bug.cgi?id=297300 Reported by: Robert Morris Reported by: syzkaller Reviewed by: vmaffione MFC after: 2 weeks Sponsored by: The FreeBSD Foundation Differential Revision: https://reviews.freebsd.org/D58678
Found with: clang -Werror=assign-enum
Found with: clang -Werror=assign-enum
No functional change, as current callers either don't check the return value or check it against OCS_HW_RTN_SUCCESS only. Found with: clang -Werror=assign-enum
DMA channels are allocated by attach_pre but released by queues_free. When iflib fails after attach_pre and before queue allocation, neither the old detach nor queues_free path releases them. Allocate channels with the TX queue state and make queues_free tolerate partially allocated rings. Use it to unwind allocation failures so TX rings are also released when RX allocation fails. An early detach can also precede PHY initialization and interrupt assignment. Skip absent PHY and channel state, and release the locks owned by attach_pre on both failure and detach. MFC after: 2 weeks
Completion queues are allocated by attach_pre but released by queues_free. An iflib failure between those stages leaks the allocation, while the original size expression also underallocates the array. Move completion queue allocation into the TX queue callback, correct its size, and unwind it with TX state if RX allocation fails. Make interrupt cleanup tolerate an unavailable array and reuse the array allocated during device initialization instead of replacing and leaking it. Release the DMA, multicast, and lock resources owned by a successful attach_pre during detach. Avoid allocating the statistics DMA area a second time near the end of attach_pre. MFC after: 2 weeks
A PF can be resetting, handling a slow link event, or deliberately withholding mailbox CTS while its VFs enumerate. Keep the VF attached when the reset handshake is temporarily unavailable so a later if_init can retry. Never leave VF hardware running without a negotiated mailbox API: start hardware only after reset succeeds, stop it when negotiation fails in attach or init, and defer later recovery through iflib. This prevents a tight reset loop while preserving recovery when the PF returns. MFC after: 2 weeks
The iflib conversion records link-status interrupts in the administrative request mask, but the administrative task did not consume them. Timer polling usually hid the omission; frequent mailbox interrupts could continually rearm that timer and leave cached link state down after hardware recovered. Claim request batches atomically, process link-setup dependencies, and sample hardware before publishing link state. Bound each invocation to eight batches and requeue residual work so a continuous producer cannot monopolize the admin taskqueue. Queue every link-related request from the legacy interrupt path. Unlike MSI-X, its threaded continuation services RX and does not enqueue the admin task. This restores the event-driven behavior of ix-3.4.39. Fixes: https://cgit.freebsd.org/src/commit/?id=b2c1e8e62049 ("ix(4): Run {mod,msf,mbx,fdir,phy}_task in if_update_admin_status") MFC after: 2 weeks
The aggregate VF mailbox poll includes only VFs whose driver configuration completed. A configured VF slot whose vf_add callback failed can nevertheless report reset, request, or acknowledgement events. Because the mailbox handler skips inactive entries, such an event remains latched and can retrigger administrative work indefinitely. Build the poll masks from every configured VF index and consume reset, message, and acknowledgement events for inactive entries without treating them as usable VFs. Use the index rather than the pool because early vf_add errors precede pool initialization. Also include E610 PFVFLREC in aggregate reset sampling. MFC after: 2 weeks
Register the vmm module handler after both the bundled device drivers and SMP. On platforms without EARLY_AP_STARTUP, SI_SUB_SMP follows SI_SUB_DRIVERS; using the later subsystem preserves the smp_rendezvous() requirement. The resulting reverse unload order performs IOMMU cleanup while every IVHD softc remains valid. Refuse an independent IVHD detach while translation state remains initialized. MFC after: 2 weeks
Do not instantiate an interrupt-remapping context for a unit whose IRTE support is disabled. In that mode the caller must retain the ordinary interrupt path. Reviewed by: kib MFC after: 2 weeks Differential Revision: https://reviews.freebsd.org/D58725
The iflib Flow Director path does not assign filters using the absolute queue and pool identifiers required by SR-IOV. Reject the combination during preflight validation rather than allowing an unsupported configuration to alter the PF receive path. The loader tunable is fixed before VFs can be created, so validation also prevents the reverse ordering of this combination. MFC after: 2 weeks
The flow_control and hdr_split variables have never been read. VF flow control is controlled by the PF, while implementing header split would require receive-path support that ixv does not provide. MFC after: 2 weeks
The shared ixgbe transmit path already creates SCTP context descriptors, and the hardware exposes the same checksum capability to VFs. Advertise it through iflib as the PF driver does. MFC after: 2 weeks
TXDCTL programming is family dependent. 82543 erratum 35 and 82544 erratum 20 require WTHRESH to remain zero; a nonzero value can corrupt descriptor writebacks and hang the controller. Leave all descriptor-control thresholds at their reset values on 82542, 82543, and 82544. On the remaining em controllers, retain the established PTHRESH=31, HTHRESH=1, WTHRESH=1, and descriptor granularity policy. Several legacy specification updates identify full descriptor writeback as a workaround for transmit descriptor-queue errata. TXDCTL bit 22 is also family dependent. It is COUNT_DESC on the 82571 family and 80003ES2LAN. Intel shared initialization explicitly sets raw bit 22 on both transmit queues of every supported ICH/PCH generation, although the integrated public documentation marks it reserved. Preserve that required setting when iflib programs the thresholds, as DPDK does. Clearing it caused a persistent I219 transmit stall under descriptor pressure. The combined em/igb setup also wrote LWTHRESH=1 on every em controller. The driver does not enable the TXD_LOW interrupt controlled by that field. Enumerate every supported em MAC type and leave the unused low-water threshold disabled. This keeps the legacy descriptor-writeback safety policies separate from igb sparse-RS operation while programming only the fields appropriate to each family. MFC after: 2 weeks
Jumbo receive tuning on integrated controllers enabled PTHRESH without a nonzero HTHRESH, contrary to the hardware programming requirements. It also covered only the integrated MAC generations present when the workaround was added. Enumerate every jumbo-capable ICH and PCH type and program PTHRESH=3 with HTHRESH=1. Linux fixed the same HTHRESH omission in b701cacdbcfb. The 82574 path combined threshold values with the reset values using bitwise OR. Requesting WTHRESH=4 while the reset value was one thus programmed five. Clear the complete threshold fields before installing the established PTHRESH=32, HTHRESH=4, WTHRESH=4 descriptor-granularity policy. MFC after: 2 weeks
iflib requests transmit completion status only on selected descriptors. Program a zero writeback threshold so igb hardware honors those sparse RS bits instead of writing back every descriptor in threshold-sized batches. Use the existing family specific prefetch threshold: eight descriptors on most controllers and 20 on I354, with a host threshold of one. These values match the Intel-derived Linux and DPDK drivers. Their nonzero writeback settings are not appropriate here because those drivers set RS on every packet. A zero writeback threshold also avoids depending on interrupt timer flushes affected by 82576 specification update erratum 26. Remove the old IGB_TX_WTHRESH macro as well. It has had no callers since the iflib conversion, so its 82575 conditional no longer implements any policy. MFC after: 2 weeks
82576 specification-update erratum 26 says MSI-X EITR expiration can fail to trigger receive descriptor writeback. A WTHRESH above one can therefore leave received packets invisible until the threshold fills. The shared threshold macros selected policy by enum ordering, so an 82576 VF fell into the generic WTHRESH=4 case. VFs always use MSI-X and require the same WTHRESH=1 workaround as the PF. Use PTHRESH=8 for 82575 and 82576 PFs and VFs, matching DPDK and the current Linux PF driver. The legacy FreeBSD PF and Linux igbvf value of 16 thrashes limited descriptor cache; no specification or erratum requires it. Retain the i354 PTHRESH=12 exception. Enumerate every supported igb PF and VF MAC type so each receives its intended policy. Also clear every threshold bit before installing the new values. The old mask retained the high WTHRESH bit, and 82575 uses six-bit fields while later controllers use five-bit fields. MFC after: 2 weeks
The transmit-ring setup was copied from the e1000 path. On I225 and I226, bits 22 through 24 are reserved and bit 25 enables the queue; it is not a legacy low-water threshold. Correct the field masks, remove the nonapplicable legacy definitions, and program only defined fields. Use PTHRESH=8 and HTHRESH=1. Keep WTHRESH at zero so the hardware honors sparse RS descriptors issued by iflib. Linux and DPDK use a writeback threshold of 16, but request status on every packet. A nonzero threshold makes hardware ignore individual RS bits and is unsuitable for the iflib completion model. The receive-ring setup likewise used a magic mask that left bit 20 of the five-bit WTHRESH field untouched. Define the receive threshold fields and replace them exactly before installing the established PTHRESH=8, HTHRESH=8, WTHRESH=4 policy. MFC after: 2 weeks
PTHRESH controls when the device prefetches transmit descriptors, HTHRESH controls how many host descriptors must be ready, and WTHRESH controls completion writeback batching. iflib places RS on selected descriptors and reclaims through those checkpoints. The data sheets require WTHRESH to be zero when software uses RS. Clear WTHRESH while retaining the established PTHRESH 32 and HTHRESH 1 fetch policy. This also follows DPDK in pairing sparse RS descriptors with WTHRESH zero. DPDK defaults to 32/0/0, while Linux ixgbevf uses 32/1/8. The 32/1/0 setting preserves FreeBSD's prefetch policy and the data-sheet requirement that HTHRESH be nonzero when PTHRESH is used. MFC after: 2 weeks
ixv uses one queue set on 82599 and X540 VFs and assumes two on X550-family VFs. The PF reports the queues assigned to each VF with GET_QUEUES after mailbox API 1.1 negotiation. Query the PF during attach. Bound symmetric iflib queue sets by the PF grant and available MSI-X data vectors. Retain one queue set per data vector: ixgbe VFs expose at most three vectors and one is reserved for the mailbox. The hardware permits each pool to use a subset of its RSS queues, so a two-queue ceiling is valid when the PF assigns four. This enables the second data vector on 82599 and X540 while avoiding an assumed second queue when an X550-family VF is granted only one. Keep the existing family limits if the mailbox is unavailable or the PF uses an older API. MFC after: 2 weeks
iflib clears IFF_DRV_RUNNING before the driver stop callback but leaves IFF_DRV_OACTIVE set. Consequently, an already queued admin task can run after the VF reset. If that task consumes a pending timer sample, it can retry failed VLAN mailbox operations and restore PF filters for the stopped VF. Continue sampling statistics, but only run the VLAN retry worker while the interface is running.
Expose the cached per-VF configuration through the iflib VF status method. Report mailbox handshake state, MAC address, access or trunk VLAN mode, hardware queue count, administrator policy, and MDD blocking state without issuing mailbox requests or reading hardware registers.
ixgbe: Report SR-IOV VF status Expose cached VF configuration, policy, and runtime state through the iflib VF status method. Include access or trunk VLAN mode, the queue count selected by the current virtualization mode, negotiated mailbox API, whether traffic is enabled, and the MDD-blocked and quarantine state. The query runs under the iflib context lock and does not issue mailbox requests or read hardware registers.
ixgbe: Add missing mailbox API 1.6 definition The SR-IOV status change reports mailbox API 1.6 but omitted its enum definition, leaving main unable to compile. API 1.6 is an established ixgbe mailbox wire revision. Add it at the end of the revision enum, before the unknown sentinel as required by the stable numbering contract. Naming the revision does not enable negotiation or operations which will come with the E610 support. Reported by: Herbert J. Skuhra <herbert@gojira.at> Fixes: https://cgit.freebsd.org/src/commit/?id=c30021fe0df9 ("ixgbe: Report SR-IOV VF status")
video: add generic video(4) capture framework Add a new video(4) framework that provides /dev/videoN, buffer management, mmap lifetime, and V4L2 ioctl dispatch for video capture drivers. Hardware drivers implement struct video_hw_ops callbacks and use video_buf_acquire/write/done to deliver frames. Reviewed by: adrian Differential Revision: https://reviews.freebsd.org/D58367
video: disable the static assertions for now The previous version of this work included the definitions but not the static asserts. It's tripping up in some CI builds, likely due to compat API building. Since this isn't any more or less broken than before, disable the static assertions until we figure out a proper path for this. Fixes: https://cgit.freebsd.org/src/commit/?id=9c9428825f4c55e3cb37412c661bb9d385db4c68 (video: add generic video(4) capture framework)
Replaced the monolithic cdevsw implementation with the video(4) framework. Reviewed by: adrian Differential Revision: https://reviews.freebsd.org/D58368
Replaced the monolithic cdevsw implementation with the video(4) framework. Reviewed by: adrian Differential Revision: https://reviews.freebsd.org/D58369
fwcam(4) no longer creates its own character device or implements the FWCAM_* ioctls; it registers with video(4) and is driven through the standard V4L2 interface on /dev/videoN. Reviewed by: adrian Differential Revision: https://reviews.freebsd.org/D58500
Raising UVIDEO_NFRAMES_MAX from 40 to 128 in 3b6f833c95eb improved
throughput on xhci but made every camera on an ehci bus fail to
stream. Integrated webcams became unusable.
Measured on a MacBookPro9,2 with two ehci(4) FaceTime HD cameras and an
xhci(4) Logitech C920:
128 32
ehci, 12 captures 0 ok 12 ok
xhci 1920x1080 5 fps 5 fps
xhci 1280x720 10 fps 10 fps
Fixes: https://cgit.freebsd.org/src/commit/?id=3b6f833c95eb
Reviewed by: bapt
Differential Revision: https://reviews.freebsd.org/D58501
Disabled the IR DMA channel on the error path, which clears the flag and frees the descriptor blocks before the chunks go away. Reviewed by: adrian Differential Revision: https://reviews.freebsd.org/D58502
Reviewed by: adrian, bapt Differential Revision: https://reviews.freebsd.org/D58503
Fix two small bugs affecting the event parsing of AMD L3 counters. AMD's manual and JSON disagree about the naming scheme on recent processors. I use the naming scheme present in the recent PPRs to be consistent, so in the JSON parser we rename 'allslices' to 'allsources' just as we already do with sliceid and sourceid. Also ensure that we parse the 0x prefix present in the newer JSON files. Reviewed by: mhorne Sponsored by: Netflix MFC after: 1 week Pull Request: https://github.com/freebsd/freebsd-src/pull/2180
A link flap left nothing in the log to work from. Both the link up and link down messages were gated on bootverbose while the message for a speed change that keeps carrier was not, so a default kernel was silent about a flap yet loud about a downshift -- the inverse of what an operator wants. The generic message from if_link_state_change() carries no speed, so gating the driver's own left the negotiated rate unrecorded. Report both transitions unconditionally. Say more than the rate. aq_hw_get_link_state() already negotiates flow control and throws it away, and Atlantic 2 reports duplex and EEE in the same link status word the rate comes from; decode them through a new get_link_info firmware op and name all of it on the up transition. EEE matters for a flap: low power idle transitions are a common source of marginal link trouble on multi-gigabit copper, and whether it was active is otherwise invisible. Give the down transition a cause. The PHY global fault code was only consulted from the thermal state machine, so an ordinary link loss reported nothing at all. Read the fault code, the firmware link state and the PHY temperature once per transition and append whatever is available. The firmware raises a fault one poll after it drops the link, so a thermal trip usually shows only its temperature here and aq_thermal_poll() names it on the following poll; the temperature alone is enough to separate a hot PHY from a cable event. Warn before the PHY trips rather than only after. The Atlantic 2 health monitor word carries a hot warning bit next to the ready and fault bits that nothing decoded. Report both edges of it from the thermal poll, so an adapter that is approaching its shutdown threshold says so while the link is still up. Expose the firmware's own link transition counters. The Atlantic 2 A0 statistics layout opens with link_up and link_down, which were read out of the firmware on every statistics poll and discarded. Publish them as dev.aq.N.fw_link_up and fw_link_down so a single flap can be told from a link that has been flapping all night. The B0 layout has no equivalent, so the op reports ENOTSUP there and the nodes are not created, matching how the temperature node is handled. Stop announcing a link state that was never read. The return value of aq_hw_get_link_state() was discarded, so a failed read would have been announced as link down. No firmware backend can fail that call today -- all three decode a register with no error path -- but the caller no longer depends on that, and it says so once if it ever starts failing. Report the hardware failures that were being discarded. The driver already reports the errors it keeps, so what stayed quiet was the set of calls whose result was never examined at all. None of these are expected to fail, which is precisely why a failure needs to say so: each one leaves the interface running but misconfigured in a way that presents as a network problem rather than a driver problem. aq_if_init() discarded aq_hw_start(), aq_hw_rss_hash_set(), aq_hw_rss_set() and aq_hw_udp_rss_enable(), so a datapath that never started or an indirection table that was never programmed showed up only as an interface that passes no traffic or delivers every flow to one queue. aq_mc_filter_apply() discarded aq_hw_mac_addr_set(), so a multicast address the stack believes is programmed could silently not be; report the address that failed and leave the filter slot for the next one instead of burning it. aq_update_vlan_filters() reported only the last of its three register writes. aq_if_stop() discarded both ring stop calls and the MAC reset, and a MAC that did not reset can still be mastering the bus. aq_if_detach() and aq_if_suspend() discarded aq_hw_deinit(). The interrupt moderation update on a link speed change was dropped as well; it runs only on a transition, so reporting it cannot become noisy. aq_if_attach_pre() discarded aq_hw_capabilities(), which is the only behavioral change here: it now fails the attach rather than continuing with an unset media type and an empty link speed mask, which would attach an interface that can never negotiate a link. It returns an error only for a device the probe table does not cover, so it is not reachable in practice. Document the resulting sysctls, along with the existing temperature and tracing nodes, which had no manual page coverage. Tested on an AQC113C (Atlantic 2 B0, firmware 1.5.38). Link up reports "speed=10000, full-duplex, flowcontrol none, EEE off", and "speed=1000" after a forced renegotiation, so the rate and duplex are read rather than assumed. A cable pull reports "link DOWN, F/W link state 0, temp 59 C" with the PHY fault clause correctly absent, which is what separates a cable event from a thermal trip. The B0 interface reports ENOTSUP for the link counters, so those two nodes are correctly not created. Traffic is unaffected: ten flows spread over all eight RX queues with no errors and no drops. Reviewed by: adrian Differential Revision: https://reviews.freebsd.org/D58749 Signed-off-by: Nick Price <nprice@FreeBSD.org>
A VF reset can sanitize its retained queue registers even when the PF does not complete the cooperative mailbox handshake. Keep those two states separate. Do not program or enable the rings until both queue sanitation and mailbox initialization have succeeded. Report either initialization failure to iflib so the interface remains stopped. Stopped admin and media-status passes now publish cached link-down state without polling the mailbox. While the VF remains administratively up, retry complete initialization after 250 ms, one second, four seconds, and then at a capped eight-second interval. Conditional iflib reset requests ensure an intervening administrative down cancels a queued retry. Avoid a redundant mailbox reset in the stop half of an immediate iflib reinitialization; the following init performs the required reset. Preserve the reset on an ordinary administrative stop and keep the existing bounded queue-sanitation retry policy independent from mailbox liveness recovery.
A failed VF reset or mailbox API negotiation currently returns from the void ifdi_init callback. Iflib then marks the interface running even though ixv left its adapter stopped. Stopped media queries can continue polling the PF, and no timer remains active to retry when the PF returns. Track mailbox readiness and report unsuccessful initialization to iflib. Stopped admin and media-status passes now publish cached link-down state without touching the mailbox. While the VF remains administratively up, retry complete initialization after 250 ms, one second, four seconds, and then at a capped eight-second interval. Preserve the requested MAC across reset, then program it once after mailbox API negotiation. The previous two pre-reset requests each could wait a full mailbox timeout after an established PF disappeared, holding the iflib context lock for about two seconds before the reset handshake. Avoid a redundant VF reset in the stop half of an immediate iflib reinitialization. Also remove the stop-time RAR mailbox request: reset has already discarded CTS at that point, and successful initialization restores the current address. Retain a reset for an ordinary administrative stop when the mailbox was established. MFC after: 2 weeks
PSRTYPE is indexed by pool in VMDq+RSS mode, and its RQPL field selects the number of receive queues available within the pool. The PF occupies the last pool, but the driver programmed pool zero and left the PF RQPL value at zero. As a result, all PF receive traffic was directed to its first queue while SR-IOV was enabled. Program PSRTYPE for the PF pool and encode its allocated receive queue count. MFC after: 2 weeks
HWRM failures currently return from the void ifdi_init callback. iflib then marks the interface running and enables interrupts despite an incomplete ring or VNIC setup. Move the hardware setup into an error-returning helper. The ifdi_init wrapper can report failure through iflib_init_failed(), while firmware recovery can propagate the same error through bnxt_open(). Also clear the initialized state after partial setup is torn down. MFC after: 2 weeks
The primary and mirror-VSI ifdi_init callbacks can return early when reset state or hardware queue and filter setup prevents initialization. Iflib then marks the interface running and enables interrupts although the driver did not finish bringing it up. Report each non-detach failure through iflib_init_failed(). Keep the existing ice reset and subinterface-reinitialization machinery responsible for scheduling recovery. MFC after: 2 weeks
The 82542-specific setup routine unconditionally reads the NVM default, overwriting a flow-control mode selected by software. It also removes transmit PAUSE support from all 82542 revisions even though the hardware restriction applies only to rev 2.0. Resolve the NVM default only when requested, scope the transmit restriction to rev 2.0, and replace integer bit masking of the enum with explicit valid mode transitions. This restores the behavior from before the Intel shared-code split and resolves -Wassign-enum. Reported by: glebius MFC after: 2 weeks
ufshci: fix data direction encoding for read commands The data_direction field in the UTP Transfer Request Descriptor is only 2 bits wide ([26:25]). UFSHCI_DATA_DIRECTION_FROM_TGT_TO_SYS was defined as 0x10, which truncates to 0b00 (No data transfer) when stored into the 2-bit field, so every read command was described to the controller as having no data phase. Only writes (0b01) happened to be encoded correctly. Define all values as 2-bit binary literals, matching the existing RESERVED = 0b11 entry, so read is encoded as 0b10 as required by the specification. Sponsored by: Samsung Electronics Reviewed by: imp (mentor) Differential Revision: https://reviews.freebsd.org/D58652
ufshci: abort submission when payload DMA mapping fails When bus_dmamap_load_mem() failed, ufshci_req_queue_prepare_prdt() manually completed and released the tracker, but its caller kept going: it built the UTRD, set the slot back to SCHEDULED, and rang the doorbell for a tracker whose request had already been freed. Return the mapping error and stop the submission so the released tracker is not resurrected. Sponsored by: Samsung Electronics Reviewed by: imp (mentor) Differential Revision: https://reviews.freebsd.org/D58653
ufshci: fail attribute reads on a non-zero config result code ufshci_uic_send_cmd() only logged the error code and returned success, so a failed DME_GET gave its caller a stale value as if it were valid. The gear and lane settings could then be programmed from that garbage. Return ENXIO for reads instead. Writes keep logging and continuing, because a device may reject an optional attribute and that must not fail bring-up. Sponsored by: Samsung Electronics Reviewed by imp (mentor) Differential Revision: https://reviews.freebsd.org/D58654
ufshci: handle controller command submit failures Return submission errors from the controller command helpers and propagate them to polled callers before waiting for completion. Free requests that never enter a hardware queue so failure paths do not leak or panic after the poll timeout. Sponsored by: Samsung Electronics Reviewed by: imp (mentor) Differential Revision: https://reviews.freebsd.org/D58655
ufshci: fix SCSI I/O request failure cleanup ufshchi_sim_scsiio() did not check the M_NOWAIT request allocation for NULL. The CDB validation and submit failure paths also returned without freeing the request. Fail the CCB when the allocation returns NULL. Free the request on every failure path. Mark the CCB as queued right before the submit, so the failure paths above do not need to touch that flag. Sponsored by: Samsung Electronics Reviewed by: imp (mentor) Differential Revision: https://reviews.freebsd.org/D58656
ufshci: free the lookup path when the periph search times out ufshci_sim_find_periph() freed the lookup path only when it found the periph. The timeout path returned without freeing it and leaked the path. Free the path at the single exit instead. Sponsored by: Samsung Electronics Reviewed by: imp (mentor) Differential Revision: https://reviews.freebsd.org/D58657
ufshci: fix WLUN periph reference counting The driver stored the WLUN periph pointer without holding a reference, so the pointer went stale when the pass(4) device went away. In addition, ufshci_sim_send_ssu() released a reference that it had never acquired. Define a simple ownership rule. ufshci_sim_find_periph() acquires the periph and returns it. The cache owns one reference. The controller destructor drops it with cam_periph_release() before taking the SIM lock, since the release takes the CAM device lock by itself. ufshci_sim_send_ssu() acquires its own reference and releases it when done. Reuse the cached periph instead of searching again, so the old reference is not leaked. Sponsored by: Samsung Electronics Reviewed by: imp (mentor) Differential Revision: https://reviews.freebsd.org/D58658
ufshci: free the correct address when DMA load fails The bus_dmamap_load() error paths passed hwq->utrd and req_queue->ucd to bus_dmamem_free(), but both pointers are only assigned after a successful load and are still NULL at that point. The freshly allocated memory was leaked. Free the local buffer instead. Sponsored by: Samsung Electronics Reviewed by: imp (mentor) Differential Revision: https://reviews.freebsd.org/D58659
ufshci: tolerate partially constructed queues in SDB teardown When attach fails, ufshci_req_sdb_destroy() runs on a partially constructed queue, and it runs twice: once from the construct error path and once from the controller destructor. Make that safe: NULL-check each resource before freeing it and clear the pointer afterwards, so a second call finds nothing to do. The construct error label no longer frees the command descriptors itself, which fixes a double free of ucd_bus_addr. Also destroy the payload DMA tag, which was previously leaked. Drop the mtx_initialized() checks: the locks are always set up before any failure path can reach the destroy. Attach can also fail before the queues were constructed at all. The destructor would then call a NULL qops.destroy pointer, so skip the destroy when the queue was never set up. Sponsored by: Samsung Electronics Reviewed by: imp (mentor) Differential Revision: https://reviews.freebsd.org/D58660
ufshci: check SDB queue allocations for failure The hardware queue and ucd_bus_addr allocations use M_NOWAIT but were used without a NULL check, and the payload bus_dmamap_create() return value was ignored, so a failed allocation was only discovered by faulting on it later. Fail the construction instead. The teardown path handles the partially constructed queue. Sponsored by: Samsung Electronics Reviewed by: imp (mentor) Differential Revision: https://reviews.freebsd.org/D58661
ufshci: do not free the devq twice on SIM attach failure cam_sim_free() with free_devq set already frees the devq, so the following cam_simq_free() call on the xpt_bus_register() and xpt_create_path() failure paths was a double free. Also clear ctrlr->ufshci_sim so a later ufshci_sim_detach() does not operate on the freed SIM. Sponsored by: Samsung Electronics Reviewed by: imp (mentor) Differential Revision: https://reviews.freebsd.org/D58662
ufshci: initialize alloc_units before the dedicated-buffer scan If every unit descriptor read failed in the LU-dedicated WriteBooster scan, alloc_units was used uninitialized. Start it at zero so that case is treated as a zero-sized buffer and WriteBooster is disabled. Sponsored by: Samsung Electronics Reviewed by: imp (mentor) Differential Revision: https://reviews.freebsd.org/D58663
ufshci: byte-swap big-endian UPIU fields The UPIU wire fields are big-endian. The task management and query builders wrote host-order values into them. The completion paths also read the results back without conversion. On a little-endian host an ABORT_TASK carried a swapped task tag and LUN, a query carried a swapped length, and attribute reads returned swapped values. Tolerant devices masked most of the damage. Convert with htobe*/be*toh at the wire boundary, as ufshci_sim.c already does for its fields. Sponsored by: Samsung Electronics Reviewed by: imp (mentor) Differential Revision: https://reviews.freebsd.org/D58664
ufshci: initialize desc_size for non-descriptor query requests The flag and attribute query builders left param.desc_size uninitialized, so stack garbage was sent as the query UPIU length field. Devices generally ignore the length for these opcodes, which hid the bug. Zero it explicitly. Sponsored by: Samsung Electronics Reviewed by: imp (mentor) Differential Revision: https://reviews.freebsd.org/D58665
ufshci: read UIC command results while holding the lock The UIC result registers (UICCMDARG2/3) are only valid between a command's completion and the next command's submission. They were read after uic_cmd_lock was dropped, so a concurrent UIC submitter could overwrite them in between. Read them into locals before releasing the lock. Also mask the generic error code to its [7:0] field when checking it, so unrelated bits in UICCMDARG2 (such as the attribute set type echoed for DME_SET) cannot be mistaken for an error. Sponsored by: Samsung Electronics Reviewed by: imp (mentor) Differential Revision: https://reviews.freebsd.org/D58667
ufshci: check completions under the queue lock The completion scan held only the recovery lock. The submit path sets a slot to SCHEDULED and then rings the doorbell, both under the queue lock. A scan running between those two steps saw a SCHEDULED slot with a clear doorbell and completed a command the device had not started. The command failed with OCS 0xf, and a reused slot could return wrong read data. Check the slot state and the doorbell under the queue lock. The submit path holds it across both steps, so a half-submitted slot can no longer be seen. Found with fio randrw verify on QEMU. Sponsored by: Samsung Electronics Reviewed by: imp (mentor) Differential Revision: https://reviews.freebsd.org/D58668
ufshci: release the CCB after sending a start stop unit command ufshci_sim_send_ssu() got a CCB from cam_periph_getccb() but never returned it. Each call leaked the CCB and one slot of the device's CCB allocation budget. When the budget runs out, the next cam_periph_getccb() waits forever and the suspend path hangs. Release the CCB while the periph lock is still held, as the other CAM periph drivers do. Sponsored by: Samsung Electronics Reviewed by: imp (mentor) Differential Revision: https://reviews.freebsd.org/D58669
ufshci: free the taskqueue on detach ufshci_ctrlr_destruct() never freed the taskqueue. Every load and unload cycle leaked the taskqueue and its kernel thread. A task that was still queued could also run after the module was gone. Free the taskqueue in destruct. Do it after the interrupt teardown so nothing enqueues new work. A reset task that is still queued at this point races the queue teardown. That race is older than this change. The planned in-flight recovery rework will close it. Sponsored by: Samsung Electronics Reviewed by: imp (mentor) Differential Revision: https://reviews.freebsd.org/D58670
ufshci: do not reset the device in the XPT_RESET_DEV handler CAM calls the SIM action callback with the SIM lock and the CAM device lock held. The XPT_RESET_DEV handler called ufshci_dev_reset(), which sleeps on device commands. Sleeping there panics when another thread contends for the lock: "panic: sleeping thread holds CAM device lock". Report success without touching the device, as nvme_sim(4) does. A real device reset needs the controller reset path. That rework is planned together with in-flight request recovery. Sponsored by: Samsung Electronics Reviewed by: imp (mentor) Differential Revision: https://reviews.freebsd.org/D58671
ufshci: return the real errno from SDB queue construction ufshci_req_sdb_cmd_desc_construct() and ufshci_req_sdb_construct() returned ENOMEM for every failure, so an EINVAL from bus_dma_tag_create() was reported as a memory shortage. Capture and return the real errno, and drop the cmd descriptor construct's now pointless out label. No functional change: no caller inspects the value beyond propagating it, so this only improves the diagnostics on an attach failure. Reviewed by: imp (mentor) Sponsored by: Samsung Electronics Differential Revision: https://reviews.freebsd.org/D58815
ufshci: pass the queue being destroyed to the cmd descriptor teardown ufshci_req_sdb_destroy() hardcoded &ctrlr->transfer_req_queue when destroying command descriptors instead of using its req_queue argument. No functional change: the branch only runs for the transfer queue, so the two pointers are always the same today. Using the argument keeps the function queue-agnostic for when more transfer queues exist (MCQ). Reviewed by: imp (mentor) Sponsored by: Samsung Electronics Differential Revision: https://reviews.freebsd.org/D58816
ufshci: validate the CDB before allocating a request The CDB pointer and length checks depend only on the CCB, so perform them before allocating and initializing the request. This avoids a wasted allocation for invalid CCBs on the I/O path and removes one request-free error path. Reviewed by: imp (mentor) Sponsored by: Samsung Electronics Differential Revision: https://reviews.freebsd.org/D58817
ufshci: consolidate the device query submit/poll pattern The six query helpers duplicated the same submit, error check, poll, and status check sequence. Move it into ufshci_dev_send_query() so future changes to the query flow are made in one place. This also unifies the failure log message format. Reviewed by: imp (mentor) Sponsored by: Samsung Electronics Differential Revision: https://reviews.freebsd.org/D58818
ufshci: correct the crypto/config register offsets and HCMID fields The reserved array after CCAP must be 508, but it was 511. This pushed the config, MCQ config, and ESI registers from 0x300 and 0x380 up to 0x900. None of these registers are used yet, so nothing broke. Also fix the HCMID bank index field. The spec places it at bits [23:16], but it was defined on top of the manufacturer code at [15:0]. Reviewed by: imp (mentor) Sponsored by: Samsung Electronics Differential Revision: https://reviews.freebsd.org/D58819
ufshci: report the highest LUN number in the path inquiry cpi->max_lun is an inclusive upper bound, but the driver reported the LUN count (8 or 32), so CAM probed one nonexistent LUN past the end. Reviewed by: imp (mentor) Sponsored by: Samsung Electronics Differential Revision: https://reviews.freebsd.org/D58820
ufshci: run the controller fail path only once Two threads could run ufshci_ctrlr_fail() at the same time. Each one walked the queues and completed the same trackers again, which caused a double free and a panic. Turn is_failed into an atomic gate, so only the first caller walks the queues. The reset task now returns early on a failed controller instead of re-enabling it. Reviewed by: imp (mentor) Sponsored by: Samsung Electronics Differential Revision: https://reviews.freebsd.org/D58944
ufshci: claim trackers before failing them ufshci_req_queue_fail() drops the queue lock to complete each tracker. In that window the completion path could complete the same tracker again. Claim the slot before dropping the lock, so the completion scan skips it. Reserved slots are left to their submit thread, which completes them itself. The manual request completion helper lost its only caller, so drop it. Reviewed by: imp (mentor) Sponsored by: Samsung Electronics Differential Revision: https://reviews.freebsd.org/D58945
ufshci: build valid fake responses for manual completion The manual completion wrote the fake response to the wrong descriptor for task management slots. It also left the task tag at zero, which tripped the task tag check under INVARIANTS. Write the fake response where the completion path reads it. Copy the task tag from the request. Reviewed by: imp (mentor) Sponsored by: Samsung Electronics Differential Revision: https://reviews.freebsd.org/D58946
ufshci: handle a recovery reset before the SIM attach When the first start attempt fails early, the recovery reset runs the start sequence again without a SIM. That pass still looked up the WLUN, so it dereferenced a NULL SIM and panicked. Attach the SIM whenever it does not exist yet. Also make the WLUN lookup return NULL when there is no SIM. Reviewed by: imp (mentor) Sponsored by: Samsung Electronics Differential Revision: https://reviews.freebsd.org/D58947
ufshci: reject new requests on a failed controller A failed controller accepted new requests, but nothing ever completed them, so the caller waited forever. The admin retry path could also resubmit a request to a dead queue. Reject new submits and admin retries on a failed controller. The submit check runs under the queue lock, so it cannot race with the queue walk in the fail path. Reviewed by: imp (mentor) Sponsored by: Samsung Electronics Differential Revision: https://reviews.freebsd.org/D58948
ufshci: fix the Snapdragon X Elite reference clock The driver's ACPI table set bRefClkFreq to 19.2 MHz. The Snapdragon X Elite feeds the device 38.4 MHz from its CXO. The firmware has no property for it. The device ran its PLL from the wrong base. Every HS mode failed. PWM still worked. The attribute is persistent. The wrong value survived reboots. Set 38.4 MHz in the table. Read the attribute first. Write it only when the value differs or the read fails. Log a changed value and a failed read. Verified on the Galaxy Book 4 Edge. Reviewed by: imp (mentor) Sponsored by: Samsung Electronics Differential Revision: https://reviews.freebsd.org/D59297
ufshci: set HS series per platform and adapt type per gear The driver always asked for Rate-B. It never set the adaptation type. The Snapdragon X Elite firmware tunes the PHY for Rate-A. A Rate-B link dies at every gear there. HS-G4 and above need initial adaptation. This is a UniPro rule. It applies to every host. Add an hs_series field to the device tables. Use Rate-A on the Snapdragon X Elite. Keep Rate-B on the PCI hosts. A table entry without an HS series fails to attach. Set PA_TxHsAdaptType to initial adaptation at HS-G4 and above. Leave it alone below that. Hosts before UniPro 1.8 do not have it. The Galaxy Book 4 Edge now links at HS-G5 Rate-A. fio results (128k sequential, 4k random, posixaio): QD | SEQ_R(MiB/s) | SEQ_W(MiB/s) | RND_R(kIOPS) | RND_W(kIOPS) ----+--------------+--------------+--------------+------------- 1 | 1357 | 1221 | 12.1 | 27.2 4 | 3103 | 3234 | 46.5 | 92.9 32 | 3508 | 3238 | 176.4 | 125.0 Sequential writes land in the WriteBooster buffer. Sustained writes drop to 556 MiB/s once the buffer runs out. Reviewed by: imp (mentor) Sponsored by: Samsung Electronics Differential Revision: https://reviews.freebsd.org/D59298
ufshci: skip the reinit when the new link works UFSHCI_QUIRK_REINIT_AFTER_MAX_GEAR_SWITCH always rebuilt the link after the gear switch. It threw away a working HS link and ended up in PWM. The reinit is only needed for a dead link. There the local side reports HS and the peer never answers. A local readback cannot tell the two apart. Peer traffic can. Probe the peer with DME_PEER_GET after the switch. Skip the reinit when the probe succeeds. Log it when the probe fails. Reviewed by: imp (mentor) Sponsored by: Samsung Electronics Differential Revision: https://reviews.freebsd.org/D59299
ufshci: tell the controller how long the EHS is The transfer request descriptor has a field for the total Extra Header Segment length. The driver left it at zero. A request that carried an EHS went out as the bare command UPIU, and the device answered a request it had only seen part of. Fill the field from the request UPIU header, which already carries the same length. Every other path sets it to zero, so nothing else changes. An EHS is the first thing that makes a request vary in size, so assert that the request and the response still fit in the command descriptor. Reviewed by: imp (mentor) Sponsored by: Samsung Electronics Differential Revision: https://reviews.freebsd.org/D59557
ufshci: add a control device node The driver only exposed a CAM SIM. Userland had no way to reach the device for anything that is not a SCSI command, so reading a descriptor or an attribute was impossible. Add /dev/ufshci%d as a root only node and the ioctl ABI header for it. The node answers no ioctl yet. The header pulls in ufshci.h, which declares bool only under _KERNEL, so include stdbool.h for userland the way nvme.h already does. Reviewed by: imp (mentor) Sponsored by: Samsung Electronics Differential Revision: https://reviews.freebsd.org/D59558
ufshci: add a passthrough ioctl This ioctl is for a port of ufs-utils: https://github.com/SanDisk-Open-Source/ufs-utils The driver only exposed a CAM SIM. Reading a descriptor, an attribute or a flag needs a query request, and a UniPro attribute needs a DME command. The driver built both only for its own setup, so userland could reach neither. Add two ioctls on the control node. UFSHCI_PASSTHROUGH_CMD sends a UPIU the caller built, sizes the request from its transaction code, and copies the response UPIU back. UFSHCI_PASSTHROUGH_UIC carries the four attribute commands and refuses the rest, which can drop the link or power the device off. It keeps the raw argument2 so the caller can read the result code the device reported, not just a failure. Validate the input and bound it by what the controller can map. The descriptor has no request length, so the controller reads it from the UPIU header, and a header that declares more than was copied in would reach past the descriptor. The ioctl layer copies output back only on a zero return. So a command that reached the device is a success even when it was refused. The caller reads the reason from the response header, and an answerless failure comes back as EIO. Clear the response before use so no stale bytes read as a device answer. Reviewed by: imp (mentor) Sponsored by: Samsung Electronics Differential Revision: https://reviews.freebsd.org/D59559
ufshci: build the ioctl file into the kernel The passthrough ioctl went into the module build only. A kernel with device ufshci then failed to link, because ufshci_ctrlr.c calls ufshci_ioctl_construct() and ufshci_ioctl_destruct() and neither was compiled in. Add the file to sys/conf/files. Fixes: https://cgit.freebsd.org/src/commit/?id=28fefc441e3b ("ufshci: add a control device node") Sponsored by: Samsung Electronics
A PF reset or loss of virtchnl service can make visible interface initialization wait up to ten seconds and then return from the void ifdi_init callback. Iflib consequently marks the interface running even though its queues were not initialized, and no retry is scheduled when the PF returns. Check reset readiness without polling during reinitialization, propagate queue-message submission errors, and bound a silent enable or disable to one mailbox timeout. Report unsuccessful initialization to iflib and publish link-down state without polling the stopped mailbox. A VFLR also discards the Admin Queue and permits the PF to replace the VF VSI. Track when full virtchnl rediscovery is required, renegotiate the API version, refresh and validate the VF resources before using a cached VSI ID, and replay the MAC and VLAN filters cleared by reset. Bound each runtime discovery attempt while preserving the existing attach-time wait. While the VF remains administratively up, retry complete initialization after 250 ms, one second, four seconds, and then at a capped eight-second interval. MFC after: 2 weeks
ixl uses head writeback by default. Hardware publishes the transmit ring head through DMA only after completing a descriptor marked RS. Marking every packet requested much more frequent head updates than iflib needs to reclaim descriptors. iflib marks selected packets with IPI_TX_INTR as completion checkpoints. It forces a checkpoint as deferred work or ring pressure grows. Retain EOP on every packet, but set RS only at those checkpoints. This batches head writebacks while preserving bounded descriptor reclamation. The optional descriptor writeback mode benefits as well. ixl already recorded only IPI_TX_INTR descriptors in its report-status queue, so status written for every other packet was not inspected. DPDK uses the same sparse RS design. Let iflib choose the adaptive interval for FreeBSD. This is a PCIe/memory bandwidth savings. MFC after: 2 weeks
iavf uses descriptor writeback by default. Hardware writes completion status into a transmit descriptor only when it completes a descriptor marked RS. iavf marked every packet RS even though its report-status queue recorded and inspected only descriptors selected by iflib. The other completion writes could not help reclaim descriptors. iflib marks selected packets with IPI_TX_INTR as completion checkpoints. It forces a checkpoint as deferred work or ring pressure grows. Retain EOP on every packet, but set RS only at those checkpoints. The deprecated head-writeback option on 700-series VFs gets the same batching: each RS checkpoint permits hardware to publish the completed ring head. DPDK uses the same sparse RS design. Let iflib choose the adaptive interval for FreeBSD. This is a PCIe/memory bandwidth savings. MFC after: 2 weeks
The VF array is zeroed at allocation, but its sysctl contexts were only populated after each VF was successfully added. If VF setup failed, IOV teardown still passed every requested VF context to sysctl_ctx_free(). An untouched context is not an initialized empty TAILQ and caused a page fault during teardown. Initialize every VF context with the array so both successful setup and partial-failure cleanup have a valid lifetime. MFC after: 2 weeks
pci_iov_enumerate_vfs() logged a failed VF creation or driver configuration but still reported the whole SR-IOV configuration as successful. The PF remained enabled with the requested NumVFs and driver state even though one or more VF children were absent. Make VF enumeration atomic. Delete children created by the failed attempt, invoke the PF driver cleanup, disable VF memory space and VF Enable, release the IOV resources, and return the original error to iovctl. Also treat failure to create a VF child as an error instead of silently accepting a partial configuration. MFC after: 2 weeks
Bound variable-length virtchnl messages before computing their expected length, following the newer Intel virtchnl implementation. Validate VF ring sizes and alignments before programming HMC contexts. DPDK uses 128-byte ring alignment and 64 through 8160 descriptors; the virtchnl ABI further specifies TX multiples of 8 and RX multiples of 32. Preserve the 4096-descriptor limit on X722. Validate queue bitmaps before changing any rings, validate all queue and interrupt contexts before applying a request, and reject invalid RSS table entries. Also avoid sending an ACK after VLAN-strip setup fails and reply to delete-VLAN errors with the correct opcode. These checks prevent malformed or oversized requests from an untrusted VF from partially programming resources outside its allocation. MFC after: 2 weeks
ixl: Make VF reset resource reconstruction fallible Treat each stage of VF reset and VSI reconstruction as fallible. Keep the VF out of VFACTIVE when PCIe drain, reset completion, VSI release, or VSI allocation fails, following the DPDK PF reset model. Propagate initial reset failures back through pci_iov_vf_add and unwind the VF queue allocation. Free the old software filter list before initializing a replacement VSI. ixl_init_filters() previously replaced the list head without freeing its entries, so every VF FLR leaked all MAC and VLAN filter objects. Reset the associated counters and VLAN bitmap with the list. Avoid allocating an initial VSI only to destroy it during the required initial VF reset, and remove redundant broadcast/filter programming from VSI setup. Also delete a partially created VSI when later Admin Queue setup fails. MFC after: 2 weeks
ixl: Enforce VF VLAN policy Add access and trunk VLAN policy to the SR-IOV schema. Access VFs use a hardware PVID and cannot alter their VLAN membership. Trunk VFs may register up to 16 VLANs, while VLAN 0 remains implicitly admitted for untagged and priority-tagged traffic. Enable hardware VLAN anti-spoofing and maintain the MAC-by-VLAN filter cross-product used by DPDK. Apply Linux's untrusted-VF limits of 18 MAC addresses and 16 VLANs so one guest cannot consume the shared PF filter table without bound. Report the effective policy through the VF status interface and document the iovctl schema. MFC after: 2 weeks Relnotes: yes
ixl: Rebuild VF resources after a PF reset A PF or EMP reset destroys the firmware switch topology, including every VF VSI. The driver rebuilt only its PF VSI and left configured VFs with stale switch element and VSI identifiers. Notify VFs before a driver initiated reset, recreate the IOV VEB, and rebuild each configured VF VSI and queue mapping after the PF switch is restored. Keep a VF out of VFACTIVE if its reconstruction fails so one failure cannot expose incomplete resources or prevent the PF and other VFs from recovering. Invalidate cached VF firmware identifiers and runtime state before recreating the VEB. If VEB creation itself fails, teardown and mailbox paths can no longer use pre-reset SEIDs or VSI data. Factor the common VEB setup out of IOV initialization so initial setup and post-reset reconstruction use the same topology and filter sequence. MFC after: 2 weeks
ixl: Report PF initialization failures to iflib ixl_if_init() returned early after AdminQ reconstruction, LAA, or VSI initialization failures. Since IFDI_INIT has no return value, iflib then marked the interface RUNNING and enabled its interrupts and timers despite the incomplete hardware state. Use iflib_init_failed() on each incomplete path. Also stop at the first ring-enable error and tear down any partially enabled rings before reporting failure. This keeps the interface stopped and makes a later initialization attempt start from a bounded state. MFC after: 2 weeks
ixl: Track and recover MDD-blocked VFs The hardware identifies each VF with TX and RX malicious-driver status latches, but the driver combined all events into one counter and reported only the last VF found. It also did not record that hardware had blocked the VF, leaving the condition invisible to management tools. Consume every PF and VF latch, keep per-direction VF counters, rate-limit per-VF diagnostics, and report the blocked and traffic-enabled state via the VF status interface. Clear the software block only after a successful VF or PF reset reconstructs its resources. Match Linux i40e policy by leaving a detected VF blocked by default. Add an opt-in hw.ixl.mdd_auto_reset_vf tunable that notifies and resets the VF for installations that prefer availability. DPDK provides the register clear and per-VF attribution precedent; Linux provides the recovery policy. MFC after: 2 weeks
ixl: Quiesce VF DMA before a PF reset A PF reset has a warning interval before the hardware reset begins. Cooperative VF drivers respond to the reset event by stopping and releasing their receive buffers, but notifying VFs did not stop the hardware queues. An active VF could therefore DMA through its old rings into freed mbuf clusters during the warning interval. Put every enabled VF in reset, drain its PCIe transactions, disable its queues, wait for receive queue shutdown, and drain transactions again before tearing down the PF HMC and AdminQ. Hold VFs in reset again while rebuilding the firmware topology. Release VF reset before programming the replacement VSI and queue mappings, since VF reset clears those registers, and publish VFACTIVE only after reconstruction succeeds. Leave a VF held in reset if rebuilding it fails. Fixes: https://cgit.freebsd.org/src/commit/?id=983e628a0c47 ("ixl: Rebuild VF resources after a PF reset") MFC after: 2 weeks
A PF link event remains cached while a VF is administratively down. Media status queries called iavf_update_link_status() and published that cached state as link-up, while the stopped admin path immediately published link-down. Consumers reacting to link events could turn this into an unbounded notification loop and prevent interface detach from draining its link-state task. Keep the cached PF state, but only publish link-up after iflib has marked the VF running. A subsequent admin pass publishes the cached state after a successful initialization. MFC after: 2 weeks
A PF reset indication leaves IAVF_STATE_RESET_PENDING set while the VF recreates its AdminQ and negotiates new resources. The ordinary AdminQ task refuses to consume messages while that state is set. Consequently, the first DISABLE_QUEUES reply after successful mailbox rediscovery remains in the receive queue and initialization times out. Later retries and manual interface restarts repeat the same cycle. Clear the stale reset indication once VERSION and GET_VF_RESOURCES have succeeded, before enabling interrupts and resuming normal virtchnl requests. MFC after: 2 weeks
This change introduces a variant of `pmap_invalidate_range` that uses the fine-grained TLB invalidation instructions introduced by the Svinval extension. These instructions allow for more efficient TLB flushing on certain implementations. Under this new scheme, `pmap_invalidate_range` was converted to an ifunc that selects the appropriate variant during boot. Event: BSDCan 2026 Reviewed by: markj, mhorne Differential Revision: https://reviews.freebsd.org/D57624
The VF link-status path can receive 2.5 and 5 Gb/s speed bits from X550-family PFs, but media reporting has no cases for them. The bootverbose message also assumes every non-10-Gb/s link is 1 Gb/s. Expose the corresponding ifmedia subtypes and derive the diagnostic speed through the shared link-speed conversion helper. MFC after: 2 weeks
X550-family devices provide clear-on-read counters for transmit and receive Low Power Idle events. Accumulate each register once in the normal statistics poll and expose the monotonic totals below the eee sysctl node. Document the counters together with the existing EEE control. Obtained from: Intel ix 3.4.39 MFC after: 2 weeks
10G-BX optics use paired wavelengths to carry 10 Gb/s Ethernet over a single strand of single-mode fiber. Their 10G compliance byte is empty, so identify them from the SFF-8472 nominal signaling rate and single-mode reach fields. When an EEPROM also advertises 1G BASE-BX10, give the complete 10G bitrate and reach signature precedence. Otherwise retain FreeBSD's permissive 1G-BX identification rather than requiring a nominal 1.3 GBd rate. MFC after: 2 weeks Relnotes: yes
The 82599 and X540 share the global RSS redirection table between the PF and its VFs. Programming that table from the PF queue count prevents a VF from using queue indices absent from the PF layout. A one-queue PF consequently directs every flow for a two- or four-queue VF to queue zero. Program at least four queue indices while SR-IOV is active. Each pool PSRTYPE.RQPL field masks the shared table to the queue subset available to that function, so the PF can continue using fewer queues. MFC after: 2 weeks
These are already available and having them defined helps keep the KASAN atomic(9) interceptors uniform. Reviewed by: mhorne MFC after: 1 week Sponsored by: The FreeBSD Foundation Differential Revision: https://reviews.freebsd.org/D58680
According to Linux 5f1c3589b0f0, the X550 PHY classifier still matches an alpha silicon ID, while the shared definitions contain the two production IDs. This can leave production hardware on the generic probing path and issue unnecessary PHY queries. MFC after: 2 weeks
The management_pkts_drpd sysctl was wired to MNGPTC, making it an alias of management_pkts_txd, instead of MNGPDC. MFC after: 3 days
Found with: clang -Werror=assign-enum Reviewed by: markj Differential Revision: https://reviews.freebsd.org/D58710
Some devices take a little longer, and the spec doesn't really seem to mandate a maximum. The common path in usbd_req_set_address() has already been bumped to 1s and I have a headset (Logitech H390) that does need a little bit longer, so let's match it in xhci. Reviewed by: aokblast Differential Revision: https://reviews.freebsd.org/D58717
Added support for E835 adapters with post-quantum cryptographic (PQC) algorithms in firmware/software signage and in SPDM attestation. Signed-off-by: Pawel Sobczyk <pawel.sobczyk@intel.com> Reviewed by: Miłosz Linkiewicz <milosz.linkiewicz@intel.com> MFC after: 1 week Differential Revision: https://reviews.freebsd.org/D57868
Add support for future client platform MFC after: 1 week
Panther Point changed the reset value of CTRL_EXT.DPG_EN to enable autonomous power gating. Clear it after hardware reset on Panther Point and Nova Point controllers to prevent unexpected Tx/Rx hangs, packet loss, or corruption. MFC after: 1 week
pmc_capture_user_callchain() checks a PMC's runcount before walking the user stack, but reads it without holding the spinlock that protects it. hardclock() can run on the same CPU during the capture and drop the runcount to zero in between, tripping the assertion and panicking INVARIANTS kernels under load. Move the check inside the existing spinlock, right where the code already confirms the sample is still valid. No functional change on kernels built without INVARIANTS. Signed-off-by: Andre Silva <andasilv@amd.com> Reviewed by: mhorne MFC after: 1 week Sponsored by: AMD Differential Revision: https://reviews.freebsd.org/D58571
pmc_capture_user_callchain() asserts that TDP_CALLCHAIN is set on the current thread, but PMC_UR samples never set that flag -- only PMC_HR and PMC_SR do. That makes the assertion always fail for PMC_UR, panicking INVARIANTS kernels as soon as pmcstat -U is used. Skip the assertion for PMC_UR. No functional change on kernels built without INVARIANTS. Signed-off-by: Andre Silva <andasilv@amd.com> Reviewed by: mhorne MFC after: 1 week Sponsored by: AMD Differential Revision: https://reviews.freebsd.org/D58572
No functional change intended. Sponsored by: The FreeBSD Foundation MFC after: 1 week
DPDK commit message net/e1000/base: fix iterator type Fix static analysis warning about comparison between types of incompatible width, which might lead to an infinite loop due to overflow. Fixes: https://cgit.freebsd.org/src/commit/?id=af75078fece3 ("first public release") Cc: stable@dpdk.org Signed-off-by: Amir Avivi <amir.avivi@intel.com> Signed-off-by: Anatoly Burakov <anatoly.burakov@intel.com> Acked-by: Bruce Richardson <bruce.richardson@intel.com> Obtained from: DPDK (3d36053991) MFC after: 2 weeks
DPDK commit message net/e1000/base: fix NVM data type in bit shift There is a static analysis warning due to wrong data types being used for NVM read data shifts. Fix it via explicit type cast. Fixes: https://cgit.freebsd.org/src/commit/?id=38db3f7f50bd ("e1000: update base driver") Cc: stable@dpdk.org Signed-off-by: Przemyslaw Ciesielski <przemyslaw.ciesielski@intel.com> Signed-off-by: Anatoly Burakov <anatoly.burakov@intel.com> Acked-by: Bruce Richardson <bruce.richardson@intel.com> Obtained from: DPDK (b932270c66) MFC after: 2 weeks
DPDK commit message net/e1000/base: fix possible variable overflow Bits can be lost as temporary math is done on signed variables and the result is assigned to an unsigned variable. Cast to u32 to force the compiler to do operations on unsigned temporary variables. Fixes: https://cgit.freebsd.org/src/commit/?id=af75078fece3 ("first public release") Cc: stable@dpdk.org Signed-off-by: Lukasz Czapnik <lukasz.czapnik@intel.com> Signed-off-by: Ciara Loftus <ciara.loftus@intel.com> Acked-by: Bruce Richardson <bruce.richardson@intel.com> Obtained from: DPDK (214cb0d7f1) MFC after: 2 weeks
DPDK commit message net/e1000/base: fix NVM loop bounds and pointer access Improve the NVM checksum routines by ensuring loop bounds are compared at the correct integer width. Use array indexing instead of explicit pointer arithmetic. Fixes: https://cgit.freebsd.org/src/commit/?id=af75078fece3 ("first public release") Cc: stable@dpdk.org Signed-off-by: Menachem Fogel <menachem.fogel@intel.com> Signed-off-by: Dima Ruinskiy <dima.ruinskiy@intel.com> Signed-off-by: Ciara Loftus <ciara.loftus@intel.com> Acked-by: Bruce Richardson <bruce.richardson@intel.com> Obtained from: DPDK (39fba42d04) MFC after: 2 weeks
DPDK commit message net/e1000/base: improve NVM checksum handling When reading NVM checksum, we may encounter the following scenarios: - Checksum may be invalid, and can be updated - Checksum may be invalid but cannot be updated because NVM is read-only For the latter case, we should just ignore invalid checksum and not attempt to update it. Signed-off-by: Sasha Neftin <sasha.neftin@intel.com> Signed-off-by: Anatoly Burakov <anatoly.burakov@intel.com> Acked-by: Bruce Richardson <bruce.richardson@intel.com> Obtained from: DPDK (5241c17f0d) MFC after: 2 weeks
Some transitional Tiger Lake systems shipped with an uninitialized checksum word. Accept that state while continuing to validate newer read-only NVM images. MFC after: 2 weeks
The shared semaphore helper accesses both the 82571 retry counter and the I210 one-time-clear flag. Those fields occupy overlapping members of the device-specific union. On 82571, incrementing the counter thus enables the I210 recovery and clears SMBI after the first timeout. Give 82571, generic 80003/82575, and I210/I211 users distinct acquire paths. Preserve the legacy peer-driver policy on 82571 and one-time recovery on I210. The separation follows the Intel e1000 base code in DPDK. MFC after: 2 weeks
DPDK commit message net/e1000/base: fix semaphore timeout value According to datasheet, software ownership of SWSM.SWESMBI bit should not exceed 100ms. Current implementation caused incorrect timeout counter values, where each iteration equals 50us delay. Because of that driver was allowed to wait for semaphore even for 1.5s. This might trigger DPC timeout. This implementation hardcodes value to 2000, which multiplied by 50us, gives 100ms of possible wait time. Fixes: https://cgit.freebsd.org/src/commit/?id=af75078fece3 ("first public release") Cc: stable@dpdk.org Signed-off-by: Pawel Malinowski <pawel.malinowski@intel.com> Signed-off-by: Anatoly Burakov <anatoly.burakov@intel.com> Acked-by: Bruce Richardson <bruce.richardson@intel.com> Obtained from: DPDK (c8bcaf0f2a) MFC after: 2 weeks
DPDK commit message net/e1000/base: fix unchecked return Static analysis has detected a write that is not checked for errors, leading to ignored error return value. Add a check. Fixes: https://cgit.freebsd.org/src/commit/?id=edcdb3c5f71b ("e1000/base: fix link flap on 82579") Cc: stable@dpdk.org Signed-off-by: Dima Ruinskiy <dima.ruinskiy@intel.com> Signed-off-by: Anatoly Burakov <anatoly.burakov@intel.com> Acked-by: Bruce Richardson <bruce.richardson@intel.com> Obtained from: DPDK (b0b6b50c20) MFC after: 2 weeks
DPDK commit message net/e1000/base: add LPI counters Add new fields in structure to indicate if EEE LPI entries have been observed on Tx and Rx path. Signed-off-by: Sasha Neftin <sasha.neftin@intel.com> Signed-off-by: Anatoly Burakov <anatoly.burakov@intel.com> Acked-by: Bruce Richardson <bruce.richardson@intel.com> Obtained from: DPDK (2e8078ee69) MFC after: 2 weeks
Accumulate the clear-on-read transmit and receive LPI event counters on EEE capable PCH and I350 family devices. Expose the 64-bit totals under the per-device eee sysctl node. MFC after: 2 weeks
sndcard_func is used as an ivar which passes around device info to the PCM and MIDI children in snd_csa(4) and snd_emu10kx(4). Simplify this and retire the need for sndcard_func, by 1) making an ivar only what used to be stored in sndcard_func->varinfo, 2) replacing sndcard_func->func with a child comparison, where needed, for instance in csa_detach(). sndcard_func is harmless in reality, but there is no reason to have the additional complexity. This way we also avoid the structure allocations. Sponsored by: The FreeBSD Foundation MFC after: 2 weeks
DPDK commit message net/e1000/base: fix reset for 82580 Fix setting device reset status bit in e1000_reset_hw_82580() function for 82580 by first reading the register value, and then setting the device reset bit. Fixes: https://cgit.freebsd.org/src/commit/?id=af75078fece3 ("first public release") Cc: stable@dpdk.org Signed-off-by: Barbara Skobiej <barbara.skobiej@intel.com> Signed-off-by: Anatoly Burakov <anatoly.burakov@intel.com> Acked-by: Bruce Richardson <bruce.richardson@intel.com> Obtained from: DPDK (88a1eb79ef) MFC after: 2 weeks
DPDK commit message net/e1000/base: fix MAC address hash bit shift In e1000_hash_mc_addr_generic() the expression: "mc_addr[4] >> 8 - bit_shift", right shifting "mc_addr[4]" shift by more than 7 bits always yields zero, so hash becomes not so different. Add initialization with bit_shift = 1, and add a loop condition to ensure bit_shift will be always in [1..8] range. Fixes: https://cgit.freebsd.org/src/commit/?id=af75078fece3 ("first public release") Cc: stable@dpdk.org Signed-off-by: Aleksandr Loktionov <aleksandr.loktionov@intel.com> Signed-off-by: Anatoly Burakov <anatoly.burakov@intel.com> Acked-by: Bruce Richardson <bruce.richardson@intel.com> Obtained from: DPDK (1749e662f6) MFC after: 2 weeks
DPDK commit message net/e1000/base: fix data type in MAC hash One of the bit shifts in MAC hash calculation triggers a static analysis warning about a potential overflow. Fix the data type to avoid this. Fixes: https://cgit.freebsd.org/src/commit/?id=af75078fece3 ("first public release") Cc: stable@dpdk.org Signed-off-by: Barbara Skobiej <barbara.skobiej@intel.com> Signed-off-by: Anatoly Burakov <anatoly.burakov@intel.com> Acked-by: Bruce Richardson <bruce.richardson@intel.com> Obtained from: DPDK (458734aaac) MFC after: 2 weeks
The multicast hash bit can be 31. Use an unsigned value so setting the bit cannot shift a signed integer into its sign bit. MFC after: 2 weeks
The i210 and i211 can occasionally fail to accept multicast table writes, particularly while addresses are added and removed rapidly. Read the table back and rewrite mismatches for up to three passes. This prevents multicast reception from retaining stale filter state while keeping the workaround limited to the affected controllers. MFC after: 2 weeks
A reset NACK from a Linux PF means that the reset completed but no permanent MAC address was assigned. Treat that response as a successful reset with a zero permanent address so attach can generate a local address instead of retrying a live mailbox. FreeBSD PFs also use a one-dword reset NACK while retained queues are being sanitized. Seed the otherwise unused request payload and accept only the three-dword, zero-filled NACK used by Linux, preserving the FreeBSD retry contract. MFC after: 2 weeks
Return a PHY write failure immediately when disabling D0 low-power link-up on 82571-family controllers. MFC after: 2 weeks
The multicast hash bit can be 31. Use an unsigned value so setting the bit cannot shift a signed integer into its sign bit. MFC after: 2 weeks
The multicast vector bit can be 31. Use an unsigned value so setting the bit cannot shift a signed integer into its sign bit. MFC after: 2 weeks
DPDK commit message net/e1000/base: fix iterator type Fix static analysis warning about comparison between types of incompatible width, which might lead to an infinite loop due to overflow. Fixes: https://cgit.freebsd.org/src/commit/?id=af75078fece3 ("first public release") Cc: stable@dpdk.org Signed-off-by: Amir Avivi <amir.avivi@intel.com> Signed-off-by: Anatoly Burakov <anatoly.burakov@intel.com> Acked-by: Bruce Richardson <bruce.richardson@intel.com> Obtained from: DPDK (3d36053991) MFC after: 2 weeks
DPDK commit message net/e1000/base: fix MAC address hash bit shift In e1000_hash_mc_addr_generic() the expression: "mc_addr[4] >> 8 - bit_shift", right shifting "mc_addr[4]" shift by more than 7 bits always yields zero, so hash becomes not so different. Add initialization with bit_shift = 1, and add a loop condition to ensure bit_shift will be always in [1..8] range. Fixes: https://cgit.freebsd.org/src/commit/?id=af75078fece3 ("first public release") Cc: stable@dpdk.org Signed-off-by: Aleksandr Loktionov <aleksandr.loktionov@intel.com> Signed-off-by: Anatoly Burakov <anatoly.burakov@intel.com> Acked-by: Bruce Richardson <bruce.richardson@intel.com> Obtained from: DPDK (1749e662f6) MFC after: 2 weeks
DPDK commit message net/e1000/base: fix data type in MAC hash One of the bit shifts in MAC hash calculation triggers a static analysis warning about a potential overflow. Fix the data type to avoid this. Fixes: https://cgit.freebsd.org/src/commit/?id=af75078fece3 ("first public release") Cc: stable@dpdk.org Signed-off-by: Barbara Skobiej <barbara.skobiej@intel.com> Signed-off-by: Anatoly Burakov <anatoly.burakov@intel.com> Acked-by: Bruce Richardson <bruce.richardson@intel.com> Obtained from: DPDK (458734aaac) MFC after: 2 weeks
Do not modify a zero-initialized PHY control value when its preceding read failed. Leave the PHY unchanged when the void power helpers cannot read its current state. This follows the defensive checks added to the corresponding e1000 helpers. MFC after: 2 weeks
A manageability VLAN can select bit 31 of its VFTA register. Use an unsigned value when constructing the register mask. MFC after: 2 weeks
The PHY capability display examines all 32 bits of the firmware bitmap. Use an unsigned value so examining bit 31 does not shift a signed integer into its sign bit. MFC after: 2 weeks
The driver already accumulates the clear-on-read transmit and receive LPI event counters. Expose the 64-bit totals under the per-device eee sysctl node. MFC after: 2 weeks
PHY identifier words are promoted to signed int when the cast is applied after the shift. Cast each 16-bit register value first so identifiers with their high bit set are assembled as unsigned data. MFC after: 2 weeks
The PHY identifier word is promoted to signed int when the cast is applied after the shift. Cast the 16-bit register value first so identifiers with their high bit set are assembled as unsigned data. MFC after: 2 weeks
The PHY identifier word is promoted to signed int when the cast is applied after the shift. Cast the 16-bit register value first so identifiers with their high bit set are assembled as unsigned data. MFC after: 2 weeks
No functional change intended. Sponsored by: The FreeBSD Foundation MFC after: 2 weeks
I225 and I226 expose three programmable LED outputs. Use LED1 for adapter identification, following the convention in DPDK. Preserve the OEM configuration across identification requests. Restore the OEM configuration before a device reset so an active led(4) pattern cannot leave the output overridden across stop or detach. The LED mode values follow the Intel I225 Software User Manual. MFC after: 2 weeks
The generic LED on and off operations do not handle internal SerDes media, leaving the led(4) device ineffective on my I210 fiber port. Use the hardware blink operation for the on phase on internal SerDes. The off phase restores the saved OEM LED configuration as before. MFC after: 2 weeks
The debug routine advanced ring pointers as if rings were contiguous. They are embedded in queue structures, so rings beyond queue zero had the wrong stride. The bogus queue index could cause an invalid MMIO read and panic the machine. Index the queue arrays first and then select the embedded ring. MFC after: 2 weeks
The debug routine reads queue registers by queue index. It also advanced unused pointers to rings embedded in queue structures. Those pointers had the wrong stride and could proceed beyond the ring object. Remove the unused pointer arithmetic. MFC after: 2 weeks
Expose the physical port identification LED through /dev/led/ix*. Save and restore the NVM-selected LEDCTL value around each request. The X550 operations also clear their PHY manual override before the register is restored. Use the dedicated firmware port-identification command on E610. Its interface selects between firmware blinking and the original mode rather than directly controlling LEDCTL. Restore the normal indication before a device stop or reset. MFC after: 2 weeks
Expose each physical port identification LED through /dev/led/ixl*. Use the existing GPIO LED helpers for most devices and the PHY provisioning interface for X710 10GBASE-T adapters. Preserve and restore the original GPIO or PHY indication mode, including before the interface is stopped. MFC after: 2 weeks
I225 and I226 report uncorrectable internal memory errors through
ICR.FER and identify the affected region in PEIND. Depending on the
region, hardware stops transmit or all PCIe and DMA traffic until the
port is reset and reinitialized.
Enable the fatal error interrupt and capture its read clear status in
the interrupt filter. Mask the cause while an iflib reset is pending,
report the affected memory regions, and expose per region indication
counters.
PCIe region parity failures require a different recovery order from a
normal reset: assert DEV_RST, wait at least 3 ms, disable PCIe master
requests, clear PCIEERRSTS, and then reinitialize the port. Follow that
sequence before entering the normal reset path and clear the remaining
LAN status afterward.
The I225/I226 PBECCSTS layout is unrelated to the PCH layout previously
copied into the igc headers. Replace those unused definitions with the
I225/I226 memory error register definitions.
Hardware validation used an I225-IT revision 3 and a debug kernel that
wrote only the documented self-clearing injection bits. It did not
synthesize interrupt or status state.
Coverage, notably DMA and Mgmt are not fully testable in my setup:
Region Observed hardware status Result
LAN PEIND 0x1, LANPERRSTS 0x200 Reset and recovered
PCIe PEIND 0x4, PCIEERRSTS 0x8 Reset and recovered
DMA DRPARC injection read back zero DFT-gated on test NIC
Mgmt Host debug strap unavailable Not injectable
The repeated LAN and PCIe tests recovered without a panic or watchdog.
A PCIe-to-LAN sequence also verified that reset-time PEIND indications
are drained before FER is unmasked.
MFC after: 2 weeks
Sponsored by: BBOX.io
I225 and I226 do not interrupt for corrected internal ECC errors.
Instead, the DMA packet buffer and PCIe memories expose sticky status
bits in PBECCSTS and PCIEECCSTS.
Sample these bits with the regular hardware statistics update, preserve
the PBECCSTS ECC enable state while clearing its RW1C indication, and
expose separate counters for the DMA packet buffer, PCIe transmit-data
memory, and PCIe retry buffer.
These counters represent observed indications rather than an exact error
count because multiple corrections between samples collapse into one
sticky status bit.
Hardware validation used an I225-IT (rev 3) and a debug kernel that
wrote only the documented self-clearing injection bits. Each test
armed the injector, exercised the owning RAM with traffic, and compared
the corresponding counter before and after.
Coverage:
Memory Observed result
DMA packet buffer corrected_dma advanced once
PCIe transmit data corrected_pcie_tx_data advanced once
PCIe retry buffer No PCIe replay source; not exercised
The retry-buffer injector requires a real PCIe replay to read the
corrupted entry. The test root port exposed AER and DPC reporting but
no protocol error injector, so ordinary traffic could not cover that
case.
MFC after: 2 weeks
Sponsored by: BBOX.io
thunderbolt: Get NHI version number from caps Sponsored by: The FreeBSD Foundation Differential Revision: https://reviews.freebsd.org/D49452
thunderbolt: Reset controllers Reset routine for both v1.0 and v2.0 routes, chosen depending on version reported in caps. Reviewed by: imp Sponsored by: The FreeBSD Foundation Differential Revision: https://reviews.freebsd.org/D49452
thunderbolt: Explicitly read NHI ISR0 register to clear it This fixes and issue where Pink Sardine controllers were not receiving interrupts for more than the first command sent on the ring. Reviewed by: emaste, imp Sponsored by: The FreeBSD Foundation Differential Revision: https://reviews.freebsd.org/D52862
thunderbolt: Support writing to router config space Reviewed by: adrian, ngie Sponsored by: The FreeBSD Foundation Differential Revision: https://reviews.freebsd.org/D49452
thunderbolt: Factor out router_prepare_cmd() Common code between router_prepare_read() & router_prepare_write(). Eventually will be used by other commands (e.g. hotplugging) aswell. Reviewed by: ngie Sponsored by: The FreeBSD Foundation Differential Revision: https://reviews.freebsd.org/D59668
thunderbolt: Change len from int to size_t in router_prepare_{read,write}()
Suggested by: ngie
Reviewed by: ngie
Sponsored by: The FreeBSD Foundation
Differential Revision: https://reviews.freebsd.org/D59668
thunderbolt: Account for CRC in router config write message size Fixes: https://cgit.freebsd.org/src/commit/?id=9c6e9bfb3474 ("thunderbolt: Support writing to router config space") Sponsored by: The FreeBSD Foundation
Query the firmware for the LEDs on each physical port and expose /dev/led/bnxt* only when alternate blinking is supported. Configure every LED in the advertised group for identification and restore its default firmware state before a function reset. This follows the DPDK and Linux bnxt HWRM identification paths. Reviewed against: DPDK, Linux Reviewed by: Sumit Saxena <sumit.saxena@broadcom.com> MFC after: 2 weeks Sponsored by: BBOX.io
Rewrite pmap_s1_invalidate_strided() to use range-based TLBI instructions when they are when available. This change can significantly reduce the number of invalidation instructions issued, leading to decreased system time. (More details on the decrease can be found in the review.) Assisted-by: Claude Code (Opus 5) Reviewed by: kib, markj MFC after: 2 weeks Differential Revision: https://reviews.freebsd.org/D58708
While testing an unrelated pmap change, D58708, that dramatically reduces the number of TLBI instructions performed, and likely the timing of unrelated events, I started seeing "Storing an invalid VFP state" panics in vfp_save_state_common(). However, the origin of this panic is elsewhere, in the else branch of sve_restore_state(). Specifically, my pmap change seems to have increased the likelihood that the thread executing the else branch would be preempted by another thread between the critical_exit() inside the else branch's call to vfp_restore_state_common() and its own call to critical_enter(). Prior to expanding the scope of the else branch's critical section, the MPASS added by this change would fire, catching the problem at its source, rather than later in vfp_save_state_common(). Assisted-by: Claude Code (Opus 5) Reviewed by: kib, markj MFC after: 1 week Differential Revision: https://reviews.freebsd.org/D58723
A PCI function can provide a host bridge into a synthetic PCI domain. Intel VMD does this: the host facing VMD function remains in its original domain while the hidden Root Ports and endpoints appear in a separate domain. The VMD function's Device Control does not describe an upstream link in that synthetic hierarchy. The hierarchy wide cold pass incorrectly used the VMD function's MPS to reprogram the hidden ports and their endpoints. Stop both cold reconciliation and runtime path walks at a PCI domain boundary. The real Root Ports within the VMD domain continue to reconcile their endpoints normally. Reviewed by: imp Tested by: Michael Butler <imb@protected-networks.net> Fixes: https://cgit.freebsd.org/src/commit/?id=8e9fe9996a1f ("pci: Reconcile MPS before attaching PCIe devices") MFC after: 6 days Sponsored by: BBOX.io Differential Revision: https://reviews.freebsd.org/D58837
Move the capability and quirk checks used by pcie_flr() into a public side effect free helper. This lets callers determine whether an FLR can be attempted before quiescing a device or saving state. The helper considers the advertised PCIe FLR capability and both the enable and disable FLR quirks. MFC after: 2 weeks Sponsored by: BBOX.io
E610 Hyper-V VFs use PCI configuration space communication instead of the native PF/VF mailbox. The generic E610 match currently attaches native mailbox operations to those devices, and the imported Hyper-V subdevice identifier is incorrect. Correct the subdevice identifier to 0x00ff, as used by DPDK shared ixgbe code, and reject that subtype until ixv has a complete Hyper-V operations table. MFC after: 1 week Sponsored by: BBOX.io
E610 VFs no longer report the actual PF link state and speed through VFLINKS. They can consequently report the default 10 Gb/s speed even when the physical link uses another rate. Negotiate mailbox API 1.6 on E610 and request the PF link state with its three-dword operation. Retain VFLINKS as the fallback when an older PF rejects API 1.6. Permit API 1.6 in the inherited xcast and queue discovery helpers so negotiating the newer revision does not disable existing operations. Use GET_QUEUES to replace E610's one-queue fallback with the grant from the PF. The common path continues to use one iflib queue set per data MSI-X vector and caps the result at two queue pairs. Preserve mailbox transport errors so the driver can distinguish an explicit PF NACK from a transient timeout. A NACK means clear-to-send state was lost and requires a VF reset. Preserve the last confirmed link state across brief transport failures and publish link down after three consecutive failures. Poll E610 link state every two seconds, matching Intel's ixgbevf service timer, and phase VFs across the intervening iflib timer ticks. This avoids a mailbox polling herd when many VFs share a PF. Media-status queries return the cached state instead of starting another synchronous exchange. An admin interrupt caused by a mailbox reply only checks for an unsolicited PF reset, preventing a request/reply interrupt loop. Hardware validation on an E610 10GBASE-T PF exercised 63 VFs. Each VF negotiated API 1.6, two queue pairs, and three MSI-X vectors. Phased polling kept 31 active VFs idle, and all 63 recovered after a PF down/up cycle without watchdogs. Adapt the API 1.6 link-state operation from DPDK shared ixgbe code. The timeout and NACK distinction follows Intel ixgbevf 5.3.25. MFC after: 2 weeks Sponsored by: Dirk-Willem van Gulik from Web Weaving (E610 hardware) Sponsored by: BBOX.io
E610 inherits the X550-family virtualization registers, anti-spoofing controls, and malicious-driver operations, but the frontend does not advertise SR-IOV and cannot negotiate the mailbox revision needed by E610 VFs. Initialize the X550-family PF/VF mailbox registers for E610 and use PFVFLREC for its VF reset events, following DPDK shared ixgbe code. Advertise the E610 SR-IOV capability, accept API 1.6 only on E610, carry the existing xcast and queue operations forward to that revision, and return the cached physical link speed and state with the three-dword E610 operation. Unsupported RSS and optional feature requests continue to receive explicit failures. SR-IOV activation also enables the existing X550-derived per-pool MDD recovery path on E610. Document the expanded protection and link-state coverage. Hardware validation created 63 VFs and rejected a 64th without flapping the running PF. Invalid TX and RX descriptor DMA independently asserted the offender's WQBR bit, gated only that VF, preserved sibling traffic, and recovered after the VF reset. FreeBSD ixv, FreeBSD DPDK, Linux ixgbevf, and Linux DPDK exercised the PF mailbox and data paths. MFC after: 2 weeks Relnotes: yes Sponsored by: Dirk-Willem van Gulik from Web Weaving (E610 hardware) Sponsored by: BBOX.io
The VF statistics registers are free running and are not cleared on read. The existing code records attach time bases and pre-reset totals, but never uses either when publishing counters. It instead replaces the low hardware bits directly, so counters can inherit pre-attach traffic or jump backward after a reset. Accumulate modular 32- and 36-bit deltas, following DPDK, while keeping the software totals across planned resets. Establish a fresh hardware baseline after each successful reset and invalidate the sampling epoch when mailbox state is lost. Detect unsolicited PF resets explicitly so a reset while link is down cannot be mistaken for counter wrap. Remove the unused base and saved-reset bookkeeping. On E610, packet and octet counters remained monotonic across a VF FLR and a PF down/up cycle. Traffic after each reset advanced both RX and TX counters. MFC after: 2 weeks Sponsored by: BBOX.io
X550 family devices provide a separate RSS key, redirection table, and MRQC register for every VMDq pool. With SR-IOV enabled, the driver continued programming only the global RSS state and never selected MRQC.MULTIPLE_RSS. VF-local RSS programming was therefore ineffective. Enable multiple-RSS mode for X550, X552, X553, and E610. Initialize the PF pool's 64-entry key, redirection table, and RSS hash controls. Leave each VF pool untouched so its driver retains ownership of its RSS key and mapping. E610 folds IPv6 extension-header traffic into its base RSS selectors and reserves the legacy EX selector bits. Translate those requested hash types rather than programming reserved bits. With two E610 VFs active and four PF queue sets, eight fixed TCP flows distributed across all four PF receive queues. E610 uses the same per-pool mode according to the E610 Datasheet, sections 7.1.3.6.2 and 8.2.2.8.20-21. MFC after: 2 weeks Sponsored by: Dirk-Willem van Gulik from Web Weaving (E610 hardware) Sponsored by: BBOX.io
PFVFRSSRK contains ten 32-bit RSS key words, numbered 0 through 9. The previous inclusive range incorrectly ended at 10. MFC after: 1 week Sponsored by: BBOX.io
e1000: Recover from PCH packet buffer ECC errors PCH LAN controllers beginning with I217 report uncorrectable packet buffer ECC errors through ICR.ECCER. Descriptor memory errors stop the MAC and require a reset before traffic can resume. Enable the interrupt on the PCH generations whose shared code setup enables packet buffer ECC. Capture the read-clear PBECCSTS value in the interrupt filter, mask ECCER while recovery is pending, and request an iflib reset from the admin task. Reenable the cause only after hardware initialization succeeds. Hardware validation used an I219-LM and the documented ICS.ECCER bit to generate the fatal interrupt. This synthesizes the interrupt cause but does not corrupt packet buffer memory or alter its ECC byte counters. Three injections in one boot each requested one reset and recovered traffic without a panic or watchdog. IMS.ECCER and PBECCSTS.ECC_ENABLE remained set after every reset. MFC after: 2 weeks Sponsored by: BBOX.io
e1000: Report PCH packet buffer ECC statistics PCH packet buffer ECC status contains read-clear byte counters for corrected and uncorrected errors. Sample them with the regular hardware statistics update and account for the snapshot captured by the fatal error interrupt path. Expose the counters and the number of reset worthy interrupt indications under dev.em.N.memory_errors. Keeping the reset counter separate also preserves evidence when another status reader wins the read-clear race. Hardware validation used an I219-LM. Three documented ICS.ECCER injections advanced fatal_resets from zero to three, exactly once per reset. corrected_packet_buffer and uncorrected_packet_buffer remained zero, as expected because ICS does not inject a memory error or alter PBECCSTS. MFC after: 2 weeks Sponsored by: BBOX.io
e1000: Recover from I210 and I211 memory errors
I210 and I211 report uncorrectable internal memory errors through
ICR.FER and identify the affected region in PEIND. Depending on the
region, hardware stops transmit or all PCIe and DMA traffic until the
port is reset and reinitialized.
Enable FER and all regional indication masks. Discard indication state
left by firmware before enabling reactions, capture the read-clear
status in the interrupt filter, and keep the cause masked while recovery
is pending. Report the affected regions and expose per-region
indication counters. Management-only errors remain under firmware
control.
PCIe region parity errors require a different recovery order from the
normal reset path. Assert the port-local CTRL.RST bit, wait at least
3 ms, verify reset completion, disable master requests, clear
PCIEERRSTS, and then enter normal port reinitialization. Do not use the
device-wide CTRL.DEV_RST sequence used by I225 and I226.
Hardware validation used an I210 revision 3 and the self-clearing
LANPERRINJ retransmit-buffer bit 9. It injected a real parity error
without synthesizing interrupt or status state.
Three injections in one boot produced the following result each time:
Observed hardware status Result
PEIND 0x1, LANPERRSTS 0x200 Reset and recovered
fatal_lan advanced exactly once per injection. All tests completed
without a panic or watchdog, and FER and the LAN parity masks remained
enabled after every recovery.
MFC after: 2 weeks
Sponsored by: BBOX.io
e1000: Report corrected I210 and I211 ECC errors I210 and I211 do not interrupt for corrected internal ECC errors. Instead, the DMA packet-buffer and PCIe memories expose sticky status bits in PBECCSTS and PCIEECCSTS. Sample these bits with the regular hardware statistics update, preserve the I210/I211 PBECCSTS enable state while clearing its RW1C indication, and expose separate counters for the DMA packet buffer, PCIe transmit data, and PCIe retry buffer. The counters represent observed indications rather than exact error counts because multiple corrections between samples collapse into one sticky status bit. Hardware validation used an I210 revision 3. Unlike I225 and I226, the published I210/I211 register definitions do not expose self-clearing injectors for these corrected ECC memories. The three counter sysctls were present and remained zero under line-rate traffic and three fatal LAN parity recoveries. PBECCSTS.ECC_ENABLE remained set after every reset. Actual corrected-error accounting was therefore not injected. MFC after: 2 weeks Sponsored by: BBOX.io
e1000: Recover from I350 memory errors I350 reports uncorrectable internal memory errors through ICR.FER and identifies the affected region in PEIND. Depending on the region and memory, hardware stops transmit, receive, or all PCIe and DMA traffic until the port is reset and reinitialized. Enable FER and all regional indication masks. Capture the read-clear status in the interrupt filter. Record the fatal PCIe, DMA, and LAN status registers, keep FER masked while recovery is pending, and expose per-region indication counters. Use the datasheet required port reset before master disable order for PCIe parity errors. Reset for PCIe, DMA, and traffic-affecting LAN errors. Statistics and VF-mailbox parity errors only require their status to be discarded and cleared; management-memory recovery remains under firmware control. Validated on an I350 (8086:1521 revision 1). Three software-set FER interrupts each advanced the unknown-region counter once, requested a single reset, restored carrier and traffic, and left FER rearmed without a watchdog. The software-set cause has no subordinate error status, so region attribution and region-specific clearing remain datasheet-based. MFC after: 2 weeks Sponsored by: BBOX.io
e1000: Report corrected I350 ECC errors I350 does not interrupt for corrected internal ECC errors. Instead, the PCIe, DMA, packet buffer, loopback, and management memories expose sticky status bits in their region-specific status registers. Sample those bits with the regular hardware statistics update, preserve the RX and TX packet buffer ECC enable state while clearing RW1C indications, and expose counters grouped by memory region. Each counter records observed indication bits rather than exact error counts because repeated corrections between samples collapse into one sticky bit. On an I350 (8086:1521 revision 1), the ECC enables remained set. All corrected-error status registers remained clear across boot, interface down/up, three FER recovery resets, and bidirectional line-rate traffic. The device has no documented corrected error injector. Therefore, the per-region paths were validated against the register definitions rather than an injected SRAM error. MFC after: 2 weeks Sponsored by: BBOX.io
e1000: Recover from 82576 memory errors 82576 reports fatal and non-fatal internal memory errors through ICR.FER and ICR.NFER and identifies the affected memory in its native PEIND layout. Fatal errors can stop transmit, receive, or both until software resets and reinitializes the port. Enable the controller-wide parity detector and implemented PEINDM reaction bits after hardware initialization, while preserving unrelated register state and omitting the absent IPsec memories on 82576NS. Enable both interrupt causes and capture the read-clear PEIND register in the interrupt filter. Keep the causes masked while the iflib admin task owns the event. Acknowledge non-fatal packet data errors without disrupting the port. Request normal port reinitialization for FER, a fatal PEIND source, or the memory hang indication. Do not apply the later I210/I350 register layout or their special PCIe parity reset order. Hardware validation used a dual-port 82576EB revision 1. Firmware left PEINDM at its 0x80000000 default; initialization explicitly programmed the parity-enable bit and produced 0xffffff07 on both ports. An NFER during two-stream TCP sustained line rate without a reset, watchdog, or carrier event. FER on the linked and disconnected ports each requested exactly one reset. The linked port resumed the existing TCP sessions after autonegotiation. PEINDM and both interrupt causes were restored after every reset. The injections set the ICR causes without corrupting SRAM, so their empty PEIND values deliberately exercised the unknown source path. MFC after: 2 weeks Sponsored by: BBOX.io
e1000: Report 82576 memory ECC errors 82576 exposes clear-on-read corrected error counters for RX, TX, switch, IPsec, descriptor-handler, PCIe retry, PCIe write, and MSI-X memories. The packet and descriptor memories also count uncorrectable errors. Sample each status register exactly once from the regular hardware statistics update and immediately before handling a memory-error interrupt. Group the counters by packet buffer, descriptor handler, and PCIe region. Skip the absent IPsec block on 82576NS. PRBESTS and PMSIXESTS are shared by both LAN ports. Attribute an indication to whichever attached port samples the clear-on-read register first so it is not counted twice. Hardware validation used an 82576EB revision 1. All nine implemented status registers reported their ECC-enable bit set. The sysctl counters remained clear across interface lifecycle, two-stream line-rate traffic, and NFER and FER cause injections. Each reset preserved the ECC enables while the driver restored PEINDM reactions. ICR cause injection does not corrupt SRAM, and the only documented data injector is specific to the IPsec packet buffer. Exact counter increments for the other memories were therefore validated against the register definitions rather than an injected ECC error. MFC after: 2 weeks Sponsored by: BBOX.io
e1000: Recover from 82575 memory errors 82575 protects its packet buffer and receive and transmit descriptor handlers with ECC. Correctable errors are repaired in hardware. Packet data errors are contained to the affected packet, while the native RX_PBUR, TX_PBUR, RX_DHER, and TX_DHER interrupt causes report unrecoverable packet buffer or descriptor handler state. The affected traffic direction remains stopped until software resets the port. Enable the three ECC blocks and hardware memory error reaction after queue and filter initialization. Capture the clear-on-read status registers in the interrupt filter and keep all four native causes masked while the iflib admin task owns the event. Request port reinitialization for every native PBUR or DHER cause. Packet data errors that do not raise a native cause remain count-only and do not disrupt the port. The captured status registers provide diagnostics and accounting but do not independently initiate recovery. Hardware validation used an 82575EB revision 2 and the documented PBEEI, RDHEEI, and TDHEEI injectors. Correctable TX/RX packet data and descriptor fetch/writeback errors preserved traffic. Uncorrectable TX/RX packet buffer header and descriptor fetch/writeback errors each requested one reset, restored traffic, and rearmed every ECC control. Repeated recovery produced no watchdogs. MFC after: 2 weeks Sponsored by: BBOX.io
e1000: Report 82575 memory ECC errors 82575 exposes clear-on-read, saturating counters for corrected and uncorrected errors in the packet buffer and the receive and transmit descriptor handlers. Sample all three status registers together from the regular hardware statistics update. When an unrecoverable event interrupts first, count the values captured by the interrupt filter so the clear-on-read status is not lost before the admin task handles it. Expose packet buffer and descriptor handler counters under the existing memory_errors sysctl node. Hardware validation used an 82575EB revision 2 and the documented PBEEI, RDHEEI, and TDHEEI injectors. Correctable and uncorrectable TX/RX packet-buffer errors and receive/transmit descriptor-handler errors advanced the corresponding counters. The controls and accounting survived repeated recovery resets and an ordinary interface down/up. MFC after: 2 weeks Sponsored by: BBOX.io
A media change can require a complete controller reset. Resetting the controller directly from the admin task leaves iflib rings, filters, and interface state programmed for the pre-reset controller. Request an iflib reset for every media change. This already was done when SR-IOV was active; use the same lifecycle for the ordinary PF case. MFC after: 2 weeks Sponsored by: BBOX.io
The reset helper discards reset_hw and init_hw errors. Runtime initialization then continues programming rings and filters, and iflib publishes the interface as running even though the controller did not reach a usable state. Initial attach similarly continues into NVM and MAC setup after a failed reset. Return errors from the reset helper. Fail attach when the controller cannot be reset or initialized, and report runtime failures through iflib_init_failed() so iflib leaves the interface stopped. Also stop register accesses and report the error when a stop-path reset fails. MFC after: 2 weeks Sponsored by: BBOX.io
The reset helper discards igc_reset_hw and igc_init_hw errors. Runtime initialization then continues programming rings and filters, and iflib publishes the interface as running even though the controller did not reach a usable state. Initial attach similarly continues into NVM and MAC setup after a failed reset. Return errors from the reset helper. Fail attach when the controller cannot be reset or initialized, and report runtime failures through iflib_init_failed() so iflib leaves the interface stopped. Also stop register accesses and report the error when a stop path reset fails. A later successful initialization completes pending fatal error cleanup and re-arms FER. Cache a requested MAC address before reset, but let init_hw program RAR0 after reset succeeds. Let iflib perform its normal attach-post failure cleanup instead of releasing the same driver resources from both layers, and make queue cleanup idempotent. MFC after: 2 weeks Sponsored by: BBOX.io
The shared code provides e1000_check_phy_82574() to recognize a PHY hang from saturated receive error and idle error counters, but em(4) never calls it. Run the check from timer driven admin work. Match Intel e1000e by requiring two consecutive positive samples before requesting a full iflib reset. MFC after: 2 weeks Sponsored by: BBOX.io
The D100 device ID was defined and handled by aq_hw_capabilities(), but had no entry in aq_vendor_info_array[], so the driver never probed it and the card was left unattached. Add the missing entry; the table lists the fibre variant last within each group, so it follows D109 rather than sorting numerically. Signed-off-by: Nick Price <nprice@FreeBSD.org> Accepted-by: adrian Approved-by: adrian (cherry picked from commit 4976b1d24abd6ff660e60041d1990b22c0bc2e5b)
aq_fw2x_thermal_arm() reached for a copper PHY register that the fibre parts do not implement, so arming failed on every init and printed a warning for a capability the hardware cannot have. Return ENOTSUP when the firmware does not advertise a temperature sensor, matching aq_fw2x_get_temp(), and warn only for a genuine failure. Signed-off-by: Nick Price <nprice@FreeBSD.org> Accepted-by: adrian Approved-by: adrian (cherry picked from commit 3c7f1aa3b831431193106f8610b2142131d774f5)
Reviewed by: mhorne Differential Revision: https://reviews.freebsd.org/D57175
This change implements the core clknode methods for the SpacemiT K1 clock control units. These methods were used to implement drivers for the APMU and PLL CCUs. The initial driver for the APMU CCU only contains clock definitions for the SDHCI controller for now. Differential Revision: https://reviews.freebsd.org/D57176 Reviewed by: mmel, mhorne, manu
Reviewed by: mhorne Differential Revision: https://reviews.freebsd.org/D57178
style(9): sys/param.h, then sys/systm.h, then the remaining kernel headers alphabetically. imgact.h belongs before proc.h. Reported by: jhibbits MFC after: 1 month Reviewed by: jhibbits Differential Revision: https://reviews.freebsd.org/D58884
T2, T1, and some pre-T1 Macs advertise a legacy PIO range in the SMC ACPI _CRS alongside a live MMIO window, but the silicon behind the PIO range is bogus. Try MMIO first, validate via LDKN >= 2, fall back to PIO if that fails or no MMIO resource is present. Drop "(T2)" from the backend message since MMIO isn't T2-exclusive. MFC: 1 week Reviewed by: ngie Differential Revision: https://reviews.freebsd.org/D58839
The problem observed on X550 adapters with 2.5 and 5 Gbps speeds negotiation on some switches is not affecting E610 adapters. Remove workaround, which omitted those speeds in the list of initially advertised speeds and advertise all speeds supported by adapter. Signed-off-by: Krzysztof Galazka <krzysztof.galazka@intel.com> Reviewed by: kbowling Tested by: Mateusz Moga <mateusz.moga@intel.com> MFC after: 1 week Sponsored by: Intel Corporation Differential Revision: https://reviews.freebsd.org/D57339
Two additional subdevice IDs were introduced to distinguish between adapters with and without manageability over USB support. Signed-off-by: Krzysztof Galazka <krzysztof.galazka@intel.com> Reviewed by: erj Tested by: Mateusz Moga <mateusz.moga@intel.com> MFC after: 1 week Sponsored by: Intel Corporation Differential Revision: https://reviews.freebsd.org/D57337
Cache PCM children and drain interrupt callbacks before detach. Allocate the parent softc by its actual size. Reviewed by: br MFC after: 1 week Differential Revision: https://reviews.freebsd.org/D58370
Expose the firmware-controlled physical port identification LED through /dev/led/ice*. Use the AdminQ port-identification command to select blinking mode and restore the netlist-selected original mode before the interface is stopped. MFC after: 2 weeks Sponsored by: BBOX.io
At two out of three call sites to vfp_restore_state_common(), the caller
must use critical_{enter,exit}() to prevent preemption between its call
to vfp_restore_state_common() and other actions, notably its call to
sve_enable(). So, it is arguably better to make
vfp_restore_state_common()'s caller responsible for performing
critical_{enter,exit}() and simply perform CRITICAL_ASSERT() inside
vfp_restore_state_common().
Reviewed by: kib, markj
MFC after: 2 weeks
Differential Revision: https://reviews.freebsd.org/D58859
Due to development history FreeBSD driver error codes are reported the same way as in Linux (as negatives) which is inconsistent with FreeBSD standard. It may cause unexpected behavior when driver errors are interpreted by a kernel as syscall handler return values. This patch converts error codes from negative to positive values for NVM access functions. Signed-off-by: Pawel Sobczyk <pawel.sobczyk@intel.com> Reviewed by: kbowling, erj, milosz.linkiewicz_intel.com Tested by: Mateusz Moga <mateusz.moga@intel.com> MFC after: 1 week Sponsored by: Intel Corporation Differential Revision: https://reviews.freebsd.org/D57642
PR: https://bugs.freebsd.org/bugzilla/show_bug.cgi?id=295453 Reviewed by: markj MFC after: 1 week Differential Revision: https://reviews.freebsd.org/D57500
bcm2835_audio_release() calls vchi_service_close() and then unconditionally calls vchi_service_release() with the same service handle. In the VCHI shim implementation, a successful vchi_service_close() calls service_free(service). The subsequent vchi_service_release() therefore dereferences a freed SHIM_SERVICE_T object when it reads service->handle, resulting in a use-after-free panic. vchi_service_release(), however, releases a reference which might block vchi_service_close() from completing successfuly, so comment it out instead of removing it altogether, until further testing is done. PR: https://bugs.freebsd.org/bugzilla/show_bug.cgi?id=297187 MFC after: 2 weeks Reviewed by: markj Differential Revision: https://reviews.freebsd.org/D58921
Add Gemini Lake LPSS SPI controller PCI IDs (0x31c2, 0x31c4, 0x31c6) to intelspi_pci_devices[]. Reviewed by: adrian Differential Revision: https://reviews.freebsd.org/D58868
On an Intel N150 host a guest started with sockets=1, cores=4, threads=1 reports "1 package(s) x 2 core(s) x 2 hardware threads" instead of four cores with one thread each, while the host itself detects its topology correctly. A FreeBSD guest picks the topology leaf in topo_probe_intel_0xb(), sys/x86/x86/mp_x86.c, and since 6badb512a94d it prefers leaf 1Fh over leaf 0Bh whenever cpu_high is 1Fh or higher. bhyve passes leaf 0 through unmodified, so the guest sees the maximum basic leaf of the host, which is 1Fh or above on Alder Lake and newer, and takes that path. x86_emulate_cpuid(), sys/amd64/vmm/x86.c, derives the topology from vm_get_topology() for leaves 1, 4 and 0Bh, but has no case for 1Fh, so the request ends up in default_leaf and the host values are returned verbatim. The guest therefore enumerates the topology of the host: with an SMT shift of 1 in the host's leaf 1Fh and four vCPUs this gives core_id_shift = 1 and pkg_id_shift = 2, which is exactly the reported 2 cores x 2 threads. Hosts whose maximum basic leaf is below 1Fh are unaffected, as the request is clamped to cpu_high before the switch statement. Leaf 1Fh uses the same level encoding as leaf 0Bh for the SMT and the core level, so handle both leaves in the same case. The module, tile and die levels are not emulated and terminate the enumeration, exactly as they already do for leaf 0Bh. PR: https://bugs.freebsd.org/bugzilla/show_bug.cgi?id=297475 MFC after: 1 week Reported by: Richard Straka <fntms@pryse.net> Reviewed by: markj Differential Revision: https://reviews.freebsd.org/D58885
Reviewed by: bz Sponsored by: NVidia networking MFC after: 1 week Differential revision: https://reviews.freebsd.org/D58941
The 82571 PBA_ECC register contains a 12-bit count of packet buffer ECC detections. The shared code enables single-bit correction, but neither FreeBSD nor the DPDK base driver consumes the counter. Sample it with the ordinary statistics timer, accumulate the value under dev.em.N.memory_errors.detected_packet_buffer, and clear the hardware counter while preserving correction and reserved register state. Do not enable its shared interrupt: the register does not distinguish corrected from uncorrectable events and does not provide a safe fatal recovery policy. Validated on a dual port 82571EB. Both functions reported zero after a clean boot, and a controlled link down/up cycle left the counter at zero while the management link recovered at 1 Gb/s without issue. MFC after: 2 weeks Sponsored by: BBOX.io
The mtw softc declares sc_epq with MTW_BULK_RX even though MTW_BULK_RX is enum value 0, while initialization and queue handling index up to MTW_EP_QUEUES; attaching a matching USB WLAN device can drive writes past the absent array and corrupt adjacent softc fields. This suggested patch sizes sc_epq with MTW_EP_QUEUES so the softc contains the endpoint queues the driver initializes and uses. Fixes: https://cgit.freebsd.org/src/commit/?id=c14b01624261 ("mt7601U: Importing if_mtw from OpenBSD") Reviewed by: bz MFC after: 1 week Differential Revision: https://reviews.freebsd.org/D58897
The rsu driver currently relies on a `KASSERT` to prove that the mbuf payload plus TX descriptor fits in the per-transfer USB TX buffer. On production kernels without `INVARIANTS`, an oversized raw 802.11 frame can reach `m_copydata()` and overwrite past that buffer, causing local kernel memory corruption. This suggested patch replaces the assertion-only guard with a runtime size check before the copy. Oversized frames return `EMSGSIZE`, leaving the existing caller cleanup paths responsible for freeing `m0`, `ni`, and the unused transfer buffer. Reachable via root / bpf access Reviewed by: bz, adrian MFC after: 1 week Differential Revision: https://reviews.freebsd.org/D58898
When 'Permit Total Port Shutdown' feature in BIOS is enabled then Port Disable bit is set in the Link Default Override Mask TLV PFA module in the NVM. In this mode, the driver acts as if the link_active_on_if_down flag is always disabled and disallow any change to that flag. This feature applies for E830 and E835 NIC series. Signed-off-by: Pawel Sobczyk <pawel.sobczyk@intel.com> Tested by: Mateusz Moga <mateusz.moga@intel.com> MFC after: 2 weeks Sponsored by: Intel Corporation Differential Revision: https://reviews.freebsd.org/D58149
The `reg` value was never initialized, so the loop could potentially
abort without a single read of the register. This was found by the
following warning from GCC:
sys/dev/thunderbolt/nhi.c: In function 'nhi_reset_v2':
sys/dev/thunderbolt/nhi.c:272:35: error: 'reg' is used uninitialized [-Werror=uninitialized]
272 | for (size_t i = 0; i < 10 && reg; i++) {
| ^~
sys/dev/thunderbolt/nhi.c:257:18: note: 'reg' was declared here
257 | uint32_t reg;
| ^~~
Reported by: GCC 15
Fixes: https://cgit.freebsd.org/src/commit/?id=efdb82413963 ("thunderbolt: Reset controllers")
chn_trigger() calls bcmchan_trigger() with the channel lock held. However, bcmchan_trigger() calls chn_intr(), which also tries to lock, which results in a lock recursion panic. chn_intr() is meant to be called by the interrupt handler and not inside CHANNEL_TRIGGER() methods. Remove the call altogether, the bcm2835_worker_play_start() call that comes after is enough. Fixes: https://cgit.freebsd.org/src/commit/?id=69cab2d1bfb5 ("Fix locking in bcm2835_audio driver") Reported by: Marco Devesas Campos <devesas.campos@gmail.com> Tested by: Marco Devesas Campos <devesas.campos@gmail.com> Sponsored by: The FreeBSD Foundation MFC after: 3 days Reviewed by: markj Differential Revision: https://reviews.freebsd.org/D59055
When the first page of a segment fits alignment, the second likely does not, so the DMA infrastructure (must_bounce()) thinks it needs to bounce the pages. Fix this by passing the previous end (address of byte following the previous segment) as a third argument to must_bounce(), so that the alignment check is done against the start of a new segment if and only if necessary, instead of the current page. Reviewed by: adrian Differential Revision: https://reviews.freebsd.org/D58627
Respect the `hw.sdhci.quirk_set` and `hw.sdhci.quirk_clear` tunables in the QorIQ eSDHC driver. This lets us tweak quirks for debugging or platform specifics.
Migrate all PowerPC QorIQ to the sdhci_fsl_fdt SDHC driver. There are a few differences that need to be accounted for: * On PowerPC device trees, the `clock-frequency` property defines the clock rate, not a `clocks` reference property. * On some older SoCs (P1022 only?) the BURST fields of the WML register are reserved, and must be 0x10, so add a FSL quirk (errata field) to account for this. * The PowerPC eSDHC controllers must have the DMA SNOOP bit set for DMA to work properly and avoid corruption. As part of this, make the fallback "fsl,esdhc" compat data work for PowerPC. If these fallbacks are not compatible with ARM SoCs, newer compat strings could be added for those, but the conservative catch-all should work for most SoCs, though perhaps less optimal. Differential Revision: https://reviews.freebsd.org/D58630
fman_qman_channel_id returns the QMan FMan channel for a given port. If a port isn't found, the wrong channel number will be returned.
* Add interrupt coalescing for DQRR and MR, with thresholds and period as tunable sysctls under the `hw.qman` tree. * Do lazy/sloppy buffer management to avoid constantly checking thresholds via QMan portal round-trips. * Add cache stashing to prewarm caches, reducing latency. * Fix the definition of Context_A in the init_fq MC command/result structures, they're 64-bit fields, not 32-bit. * Reorder the dpaa_eth_frame_info as a bit of cleanup. * Take advantage of the fact that UMA small allocations are returned in the DMAP, and avoid pmap_kextract(). These changes together improve throughput by ~1.5% (925Mbps->935-940Mbps) consistently, and reduce CPU usage by a bit, increasing idle CPU from 30%->35% minimum.
Set the verb correctly so the query works. Also print out the programmable FQID fields as well.
Apply 6464974 to dTSEC, since it supports the same offload capabilities as mEMAC. The DPAA_CSUM_TX_OFFLOAD macro moves from if_memac.c to the shared dpaa_eth.h since both drivers now reference it.
uaudio20_set_speed() split the sample rate into bytes by hand. Use uDWord and USETDW() instead. No functional change intended. Sponsored by: The FreeBSD Foundation MFC after: 1 week Reviewed by: markj Differential Revision: https://reviews.freebsd.org/D59066
Fixes: https://cgit.freebsd.org/src/commit/?id=550db3d6f502 ("sdhci: Initial support for the SpacemiT K1 sdhci controller") Reviewed by: bnovkov Sponsored by: The FreeBSD Foundation Differential Revision: https://reviews.freebsd.org/D59083
Any failure within bus_alloc_resources() will call bus_release_resources(); thus the call is redundant here. MFC after: 3 days Sponsored by: The FreeBSD Foundation
Fix reporting of state and capabilities by the gpioctl command. Support selection of pull-up and pull-down resistors. Support second gpio device (AON - always on power domain) to allow attaching gpioled device to visionfive2 status LED or querying boot selection switches. Reviewed by: mhorne MFC after: 1 week Differential Revision: https://reviews.freebsd.org/D58693
intelspi: sort PCI ID table by device ID Sort the existing LPSS peripheral SPI controller PCI ID table by numeric device ID so new entries have an unambiguous insertion point. Reviewed by: wulf Differential Revision: https://reviews.freebsd.org/D58996
intelspi: add Broxton SPI controller IDs Add PCI device IDs for Broxton-generation LPSS peripheral SPI controllers. Reviewed by: adrian Differential Revision: https://reviews.freebsd.org/D58997
intelspi: add Apollo Lake SPI controller IDs Add PCI device IDs for Apollo Lake-generation LPSS peripheral SPI controllers. Reviewed by: adrian Differential Revision: https://reviews.freebsd.org/D58998
intelspi: add Cannon Lake SPI controller IDs Add PCI device IDs for Cannon Lake-generation LPSS peripheral SPI controllers. Reviewed by: adrian Differential Revision: https://reviews.freebsd.org/D58999
intelspi: add Comet Lake SPI controller IDs Add PCI device IDs for Comet Lake-generation LPSS peripheral SPI controllers. Reviewed by: adrian Differential Revision: https://reviews.freebsd.org/D59000
intelspi: add Wildcat Lake SPI controller IDs Add PCI device IDs for Wildcat Lake-generation LPSS peripheral SPI controllers. Reviewed by: adrian Differential Revision: https://reviews.freebsd.org/D59011
intelspi: add Panther Lake SPI controller IDs Add PCI device IDs for Panther Lake-generation LPSS peripheral SPI controllers. Reviewed by: adrian Differential Revision: https://reviews.freebsd.org/D59012
intelspi: add Nova Lake SPI controller IDs Add PCI device IDs for Nova Lake-generation LPSS peripheral SPI controllers. Reviewed by: adrian Differential Revision: https://reviews.freebsd.org/D59013
ofw_pcibus: Inherit PF locality for SR-IOV VFs PCI VFs are allocated dynamically and have no corresponding OFW node. The zero-filled OFW PCI devinfo currently leaves obd_node as 0, which is not the invalid-node sentinel and can send NUMA lookup through an unrelated firmware node. Initialize dynamically allocated devinfo with an invalid OFW node. For VF locality queries, use the owning PF's node when it exists. Fall back to the PCI bus when neither the VF nor PF has a firmware node. This preserves existing CPU-locality behavior for ordinary PCI devices while making VF domain and interrupt placement follow their PF. Reviewed by: PowerPC (jhibbits) MFC after: 2 weeks Sponsored by: BBOX.io Differential Revision: https://reviews.freebsd.org/D59064
pci: Expose a VF's owning PF to bus subclasses ofw_pcibus now uses pci_iov_get_pf() to inherit PF locality for VFs, but the accessor was inadvertently left in an uncommited ACPI change. This breaks powerpc builds. Expose the accessor from the PCI core and provide a stub when PCI_IOV is omitted. Record VF ownership before pci_add_child() so child added callbacks can safely query it, and remove the later redundant assignment. Fixes: https://cgit.freebsd.org/src/commit/?id=f003e86335c9 ofw_pcibus: Inherit PF locality for SR-IOV VFs MFC after: 2 weeks Sponsored by: BBOX.io
BUS_GET_DOMAIN can report a PCI function's firmware locality, including an SR-IOV VF's inherited PF locality, but ordinary OFW PCI functions still use the shared bus DMA tag. Consequently, busdma metadata and coherent memory can be allocated from the bus's domain instead of the function's domain. Create and cache a child tag for each function that requests a DMA tag and apply its reported domain without modifying the shared parent tag. Apply the same domain to the private IOMMU tag already created by the pSeries PCI bus. Destroy cached tags when PCI children are removed so VF create and destroy cycles do not leak them. Reviewed by: PowerPC (jhibbits) MFC after: 2 weeks Sponsored by: BBOX.io Differential Revision: https://reviews.freebsd.org/D59065
Cache PCM children instead of calling device_get_children() from the interrupt handler. Drain callbacks before child detach so cached pointers cannot outlive the PCM softc. Allocate the parent softc by its actual size. This mirrors snd_hdspe's interrupt dispatch and detach lifecycle. Reported by: christos MFC after: 1 week
The TSO workaround splits the final DMA segment to create a four byte sentinel descriptor. Intel documents the premature descriptor writeback erratum and this workaround in the 82540EP and 82545GM specification updates (erratum 3) and the 82546GB specification update (erratum 1). Limit the workaround and its preceding TSO state to the legacy PCI and PCI-X controllers so PCIe controllers retain their natural descriptor layout using one fewer descriptor per TSO packet, no split of the final segment, and one less four byte DMA. MFC after: 2 weeks Sponsored by: BBOX.io
intelspi: add Lakefield SPI controller IDs Add PCI device IDs for Lakefield-generation LPSS peripheral SPI controllers. Reviewed by: adrian Differential Revision: https://reviews.freebsd.org/D59001
intelspi: add Ice Lake SPI controller IDs Add PCI device IDs for Ice Lake-generation LPSS peripheral SPI controllers. Reviewed by: adrian Differential Revision: https://reviews.freebsd.org/D59002
intelspi: add Jasper Lake SPI controller IDs Add PCI device IDs for Jasper Lake-generation LPSS peripheral SPI controllers. Reviewed by: adrian Differential Revision: https://reviews.freebsd.org/D59003
intelspi: add Tiger Lake SPI controller IDs Add PCI device IDs for Tiger Lake-generation LPSS peripheral SPI controllers. Reviewed by: adrian Differential Revision: https://reviews.freebsd.org/D59004
intelspi: add Elkhart Lake SPI controller IDs Add PCI device IDs for Elkhart Lake-generation LPSS peripheral SPI controllers. Reviewed by: adrian Differential Revision: https://reviews.freebsd.org/D59005
intelspi: add Alder Lake SPI controller IDs Add PCI device IDs for Alder Lake-generation LPSS peripheral SPI controllers. Reviewed by: adrian Differential Revision: https://reviews.freebsd.org/D59006
intelspi: add Raptor Lake SPI controller IDs Add PCI device IDs for Raptor Lake-generation LPSS peripheral SPI controllers. Reviewed by: adrian Differential Revision: https://reviews.freebsd.org/D59007
intelspi: add Meteor Lake SPI controller IDs Add PCI device IDs for Meteor Lake-generation LPSS peripheral SPI controllers. Reviewed by: adrian Differential Revision: https://reviews.freebsd.org/D59008
intelspi: add Arrow Lake SPI controller IDs Add PCI device IDs for Arrow Lake-generation LPSS peripheral SPI controllers. Reviewed by: adrian Differential Revision: https://reviews.freebsd.org/D59009
intelspi: add Lunar Lake SPI controller IDs Add PCI device IDs for Lunar Lake-generation LPSS peripheral SPI controllers. Reviewed by: adrian Differential Revision: https://reviews.freebsd.org/D59010
Perform the allocations outside the lock section so that we can use M_WAITOK. Holding the lock here is actually not really necessary and we could just as well remove it, but keep it for consistency. Sponsored by: The FreeBSD Foundation MFC after: 1 month Reviewed by: kib Differential Revision: https://reviews.freebsd.org/D59079
The TX SG-build loop in dpaa_eth_if_start_locked() walked page boundaries with PAGE_MASK arithmetic even for buffers that lived entirely within one page -- the common case, since MCLBYTES is smaller than PAGE_SIZE. Add a fast path that emits a single SGT entry for wholly-in-one-page segments and skips the inner while entirely. Fix the following bugs while we're here: 1. "if (m->m_len == 0) continue;" in the outer loop never advanced m -- any zero-length mbuf hung the TX path in an infinite loop. Fix this by switching to a for loop, with the advancement in the post-clause. 2. In the inner (page-splitting) loop, the cap "if (m->m_len < ssize) ssize = m->m_len;" compared against the mbuf's original length, not the remaining bytes. A single mbuf whose data started mid-page and ran into a second page would produce a second SGT entry with ssize > rem, over-reading past the buffer end into whatever followed in kernel memory. Fixed by tracking a local rem and capping ssize against it. 3. If the whole mbuf chain consisted of zero-length segments, the final-flag store "fi_sgt[i - 1].final = 1" wrote to index -1. Reject empty frames up front now instead. As part of this, rename the inner counter from dsize to rem for clarity instead of playing double-duty in both inner and outer loops.
* Set qman_channel_base after determining if QMan is v3, otherwise this global stays at 0x21, which messes up the shift in qman_portal_static_dequeue_channel(). * Fix the base shift in qman_portal_static_dequeue_channel(), there are only 15 channels available, not 16, so starting at a shift of 15 yields shifting into the portal-specific channel. * Correct vmem pool names for QMan resource pools.
82580 reports fatal parity and uncorrectable ECC errors through ICR.FER and its four region PEIND hierarchy. Region specific status registers identify PCIe, DMA transmit, DMA receive, DMA host, and LAN port memories that can leave traffic stopped. Enable the documented DMA, PCIe, packet-buffer, and host-owned LAN parity and ECC checks only after initializing queue and filter tables. Leave the flexible filter parity controls under management firmware ownership. Capture read-clear and RW1C status in the interrupt filter and keep FER masked until the admin task resolves the event. Reset for a host-owned region or an unknown FER source. Leave management-only recovery to firmware. Use CTRL.RST before master disable because fatal 82580 memory errors can stop PCIe traffic. Do not use CTRL.DEV_RST: specification update item 9 declares that bit reserved and says it must always be written as zero. Wait for EEPROM auto read completion; STATUS bit 21 is reserved on 82580, not PF_RST_DONE. Validated on an Intel I340-T2 (82580, revision 1). A one queue port programmed LANPERRCTL as 0x6e00 while a four-queue port programmed 0x7e00, avoiding the RSS checker until RETA is initialized. Both ports programmed PEINDM 0xf, DTPARC 0x1555, DRPARC and DDPARC 0x55, PCIEERRCTL 0x5555, and PCIEECCCTL 0x11. Three down/up cycles left all status and counters clear. A synthetic ICS.FER event on each function caused exactly one unknown source reset without advancing the sibling function counters. Controls were restored, the linked port recovered carrier, and bidirectional traffic after recovery. Enabling flexible filter parity checkers before programming their memories produced genuine LAN region faults with LANPERRSTS bits 0 and 1. Each fault advanced fatal_lan and fatal_resets exactly once, left the sibling function unchanged, and recovered the port proving hardware events will trigger the intended recovery. MFC after: 2 weeks Sponsored by: BBOX.io
82580 exposes clear-on-read, saturating corrected error counters for the receive and transmit packet buffers. Its two PCIe command memories expose RW1C indications for uncorrectable ECC errors. Sample the packet buffer counters and PCIe indications from the regular hardware statistics update. Fatal recovery samples the PCIe indications from the serialized admin path rather than the interrupt filter. Thus, either the regular statistics pass or recovery reads and clears each indication, but they cannot both account it. Also preserve indications observed while initialization is completing. Expose the exact packet buffer error total and observed PCIe command memory indications under the memory_errors sysctl node. Multiple PCIe errors between samples can collapse into one indication per memory. Validated on an Intel I340-T2 (82580, revision 1). A clean boot and three down/up cycles left the packet-buffer, PCIe, and region-specific counters at zero. Synthetic ICS.FER events advanced fatal_unknown and fatal_resets exactly once on the targeted function without changing the sibling or ECC counters. The 82580 datasheet exposes no ECC or parity error injection register, so corrected packet buffer and PCIe ECC accounting could not be forced independently. MFC after: 2 weeks Sponsored by: BBOX.io
The shared base code already selects and configures the 82598 BX, 82599 KR, 82599 SFP Express Module, X552 XFI, X553 QSFP, and X553 N QSFP device IDs, but the FreeBSD probe table omits them while DPDK lists them. MFC after: 2 weeks Sponsored by: BBOX.io
The RAPL probe read MSR_RAPL_POWER_UNIT with a bare rdmsr(). RAPL is not enumerated by CPUID on either vendor and the register is absent on older Intel and AMD parts and under a hypervisor that does not emulate it, so the read raises #GP and loading hwpmc panics the machine. Read it with rdmsr_safe() and return ENXIO when it is not there, as this function already does for the energy MSRs. Both callers already drop the class when the probe fails. Fixes: https://cgit.freebsd.org/src/commit/?id=a99d04f39dab ("hwpmc: add RAPL energy-counter class (AMD + Intel)") Assisted-by: Claude Code (Opus 5)
An energy status unit of zero means one joule per raw tick, which no part reports; it is what a hypervisor returns for an MSR it does not implement. Both energy rows are scaled by that field, so the class would be registered with counters that read zero forever. Refuse it, as the class is already refused when no energy MSR responds. Assisted-by: Claude Code (Opus 5)
Validate the IVRS table and every subtable length before using either to form iterator bounds. Reject truncated typed IVHD blocks instead of passing them to a type-specific callback. Within each IVHD payload, correct the lower-bound comparison for extended range entries and validate fixed-size entries, paired range terminators, the fixed HID body, and the variable HID UID before dereferencing or advancing. Malformed firmware can no longer drive either iterator beyond its enclosing object. Reviewed by: kib MFC after: 2 weeks Sponsored by: BBOX.io Differential Revision: https://reviews.freebsd.org/D58724
This would be used to switch to the "WiFi bank" before reading the rom. The 8723bu will need this, currently a nop on all chips. Differential Revision: https://reviews.freebsd.org/D59106
The endpoints we want won't always be on interface 0. Instead, allow the interface index to be specified in driver_info when probing. Reviewed by: adrian Differential Revision: https://reviews.freebsd.org/D59107
Add plumbing for future FMan KeyGen-driven multi-queue RX. Pure infrastructure; no behavioural change for existing single-FQ consumers. qman: * New qman_percpu_channel(cpu) to get the per-CPU channel, needed for receive-side scaling. * New qman_alloc_fqid_range(count, *basep) / qman_free_fqid_range() reserve a contiguous FQID range so a later KeyGen-distribution caller can compute FQID = base + (hash & mask) and create each FQ individually with force_fqid=true (each landing on its own per-CPU channel). qman_fq_create: * Honor the force_fqid / fqid_or_align parameters: when force_fqid is set, use the caller-supplied FQID and skip the internal vmem_alloc. The fqids_num != 1 restriction is lifted; qman_fq_list[] now records the handle at every FQID slot in the range so DQRR dispatch works for the whole range. * Add fqid_count and force_fqid to struct qman_fq so qman_fq_free() can retire every FQ in the range.
Preparation for FMan KeyGen-driven multi-queue RX. Replace the single sc_rx_fq / sc_rx_fqid pair with a sc_rx_fqs[] array (currently one entry) and sc_rx_fqid_base. Each entry carries a back-pointer to the softc for use by the RX callback.
Add sys/dev/dpaa/fman_keygen.[ch]. Public API is four functions:
* fman_kg_init(sc) -- Initialize KeyGen subsystem, clear out any stale
config.
* fman_kg_fini(sc) -- Teardown KeyGen
* fman_kg_alloc_hash_scheme(sc, port, base_fqid, nfqs)
-- Allocate a scheme, program it for
RSS-over-IP-5-tuple hashing to nfqs FQs
starting at base_fqid, bind it to port.
* fman_kg_free_hash_scheme(sc, port)
-- Remove a scheme added by
fman_kg_alloc_hash_scheme().
KeyGen state (bitmap + port->scheme table) is added to the fman softc.
Future work may allow configuring the KG hash inputs, but what we have
now (5-tuple of src/src-port/dst/dst-port/IPSec SPI field) is
sufficient.
Grow sc_nrxfqs from 1 to the CPU total, and hash the RX 5-tuple across the range with the KG driver from the prior commit. Each FQ lands on its own per-CPU QMan channel, so a given core drains only its own share of RX work and gets frame annotation + data-head stashed into its cache. * Add alignment parameter to qman_alloc_fqid_range() to meet KeyGen requirements. * Initialize 1 frame queue (FQ) per CPU in dpaa_eth_fm_port_rx_init(), using a 5-tuple to spread the load across CPUs. * Channel ownership for TX confirms moved from rx_init/free to tx_init/free -- sc_rx_channel is now a TX-confirm-only per-port pool channel. Fallbacks/degradation: * If any per-CPU channel is -1 (no portal attached) the port fails to attach with a clear message. * If the FQID range can't be allocated aligned, the port fails attach. * If KG scheme allocation fails at port setup, the port keeps its N FQs but only FQ #0 sees traffic.
Reduce the code executed in the DQRR dequeue loop, and move the heavy-weight operations to post-dequeue loop. * Batch if_input() after DQRR dispatch loop completes. Only do the DQRR_CI_CINH write at the end of the loop, so only up to 16 entries will be processed. * Add software LRO per FQ. Each per-CPU RX FQ gets its own LRO tracking structure. Since LRO is configured at FQ initialization time, allocate the ifnet earlier in attach to prevent a panic.
The 82579 PCIm2PCI arbiter can acknowledge a host MAC CSR write while the Management Engine is accessing another CSR. The host write can be lost; subsequent target accesses may no longer be claimed by the MAC and can hang the system. For 82579 controllers with valid management firmware, wait for the ME CSR access indication before every MAC CSR write. Keep the wait bounded and use DELAY because writes occur in interrupt and datapath contexts. Verify every transmit and receive tail write. If a tail does not hold the requested value, disable its datapath direction and request a full iflib reset. Keep the ordinary register-write path as a direct MMIO write behind a predicted per-device gate. Contain the wait and tail recovery in the 82579 slow path rather than adding tail-specific accessors and state to the rest of the e1000 family. Documentation on the PCH NICs is scare so Intel's Linux e1000e fixes publicly document the hardware failure and required serialization as commits bdc125f73f3c and d601afcae2fe. This implementation is a bit cleaner. Tested on a Thinkpad T430 (82579LM) with a test kernel to simulate ME contention without incident as well as lost tail writes causing a succesful recovery. MFC after: 2 weeks Sponsored by: BBOX.io
When building struct arm64_bootparams to pass to initarm we store the boot_el field as a 64-bit value, however it is defined as an int which is 32-bits. Switch to store using the 'w' register as this will store a correctly sized value. Previously this would trash the trailing padding, so is only a correctness issue. Reviewed by: emaste Sponsored by: Arm Ltd Differential Revision: https://reviews.freebsd.org/D58988
Co-developed-by: Andrew Turner <andrew@FreeBSD.org> Sponsored by: Arm Ltd
Sponsored by: Arm Ltd
Fix the non-VHE register field definition of CNTHCTL_EL1PCTEN to use the correct non-VHE shift. Signed-off-by: Kajetan Puchalski <kajetan.puchalski@arm.com> Sponsored by: Arm Ltd Pull Request: https://github.com/freebsd/freebsd-src/pull/2378
Serialize S3X I/O and cap dtransfers while keeping namespace handling. Select 64/128-byte submission queue entries explicitly and set CC.IOSQES from the same value used for the software queue stride. When fatal status is set, wait for pending PCIe transactions and then force FLR so a wedged controller doesn't panic or timeout. MFC: 1 week PR: https://bugs.freebsd.org/bugzilla/show_bug.cgi?id=296946 Fixes: https://cgit.freebsd.org/src/commit/?id=5e0ba47aa00e Reviewed by: ngie, imp Differential Revision: https://reviews.freebsd.org/D58821
The Apple S3X controller exposes internal namespaces beyond NSID 1 that aren't meant to be visible to the OS. Added QUIRK_APPLE_S3X_NS1_ONLY and nvme_ctrlr_num_namespaces()/nvme_ctrlr_nsid_visible() helpers, and route namespace construction, notification, and AER namespace-changed handling through them instead of a raw cdata.nn count. MFC after: 1 week Reviewed by: imp Differential Revision: https://reviews.freebsd.org/D58844
There are several port functions that are only for one type or the other, so they don't make sense to be together. Splitting these up also simplifies adding support for the Offline/Host Command ports.
Bits [0:1] map to "B", the "Segment size selector", and are not part of the VSID. Correct this mask.
Using hex here breaks the instruction generated by MRS_REG_ALT_NAME. Switch to a decimal value. Sponsored by: Arm Ltd
Some ID_AA64ISAR2_EL1 fields values are incorrectly indented. Values have an extra space before the macro to make scanning for them easier. Add this extra space to the two fields that were missing it. Sponsored by: Arm Ltd
The C standard does not require diagnostic messages to be quoted, but some tools get confused by unbalanced quotes such as the apostrophe in “don't”. Wrap this message in double quotes to resolve the confusion. Sponsored by: Klara, Inc. Sponsored by: NetApp, Inc.
This function has a loop where it attempts to lock all channels in a group. If doing so would block, it releases all locks, sleeps for a bit, and tries again. However, once the syncgroup lock is dropped, nothing prevents the syncgroup structure from being freed. Fix the inner loop: after waking up, break out of it unconditionally and start everything again. I think the old code was also buggy and not well-exercised: after waking up we'd continue to try and continue locking channels. Then we'd try again from the beginning and fail to lock the channels we had already locked. Approved by: so Security: FreeBSD-SA-26:58.sound Security: CVE-2026-58091 Reported by: Hazley Samsudin of GovTech CSG Reviewed by: christos Sponsored by: The FreeBSD Foundation Differential Revision: https://reviews.freebsd.org/D58912
Move vgic_v3 structures in preparation for vgic interface rework for GICv5 support. Reviewed by: Sarah Walker <sarah.walker2@arm.com> Sponsored by: Arm Ltd
Reviewed by: Sarah Walker <sarah.walker2@arm.com> Sponsored by: Arm Ltd
"1 << 31" is a signed int left shift 31 which is undefined as the shift is too large. Use an unsigned int to make the value defined. Sponsored by: Arm Ltd
The Logitech H390, for instance, has the following interface layout:
~~
7 INPUT 34 INPUT 10 INPUT
Mic (0x201) Mic (0x201) USB Stream (0x101)
| | |
v v |
19 FEATURE 35 FEATURE |
| | |
v v |
25 EXTENSION +------> 36 MIXER <---+
| |
v v
13 OUTPUT 22 FEATURE
USB Stream (0x101) |
v
16 OUTPUT
Speaker (0x301)
~~
The 7->13 path on the left is a typical microphone-in configuration,
while the right side is a little more complicated. The 34 -> 35 -> 36
leg is describing a hardware sidetone control, while the other is a
standard audio-out configuration.
During feature unit evaluation, we need to pick up the scenario of node
35 above, which is directly wiring the microphone to the speaker. Right
now we'll likely tie it to the pcm/vol levels and this unit will emit
a very prompt feedback screech, but it's really shaped more like a
MONITOR control.
This avoids mishandling feature unit 22 because that's evaluated in one
of the other cases: one of the inputs is the USB stream, so it's
wired up as a PCM.
One note on this headset: the presence of mixer 36 currently breaks the
`vol` control, leaving only `pcm` to control the volume. Given that it
has both Mic and USB input, I suspect we get a 1:1 cluster configuration
for the Mic input but something more complicated for the USB input that
we end up ignoring. Thus, "vol" might technically control the monitor
volume but isn't wired up to the USB input cluster. I have not had a
chance to confirm this, yet.
PR: https://bugs.freebsd.org/bugzilla/show_bug.cgi?id=291424
Reviewed by: christos
Differential Revision: https://reviews.freebsd.org/D58779
The previous patch assumes that we don't want to reset toggle bit in STOPPED_STEP. However, a device can explicitly call usbd_clear_data_toggle if necessary. As a result, instead of not dropping the bit unconditionally, we added a field in xhci to specify that we want to drop it, so that usbd_clear_data_toggle can handle it correctly. Reported by: oh Reviewed by: kevans Tested by: oh Fixes: https://cgit.freebsd.org/src/commit/?id=28d85db46b48 ("xhci: Do not drop and add bits in xhci") MFC after: 3 days Differential Revision: https://reviews.freebsd.org/D59186
SR-IOV VFs are instantiated dynamically from their PF rather than enumerated from ACPI. A VF's runtime slot and function can match an unrelated _ADR below the bridge. acpi_pci_save_handle() stores that handle in the VF's devinfo before acpi_pci_update_device() runs. If the handle is already bound to another device_t whose parent is not acpi0, acpi_pci_update_device() panics under INVARIANTS. Without INVARIANTS, the VF retains the unrelated handle, so subsequent ACPI lookups, including NUMA and power-management operations, can act on the wrong namespace node. Skip ACPI namespace matching for VFs. Reviewed by: jhb MFC after: 2 weeks Sponsored by: BBOX.io Differential Revision: https://reviews.freebsd.org/D59061
SR-IOV VFs are instantiated from their PF and intentionally do not receive an ACPI handle by matching their runtime BDF. Consequently, ACPI locality queries for a VF fall back to the upstream bus. This is usually sufficient, but loses a _PXM supplied specifically for the PF. Use the PCI core's owning-PF accessor for BUS_GET_DOMAIN and BUS_GET_CPUS requests made for a VF. This preserves the VF's lack of an ACPI handle while allowing its CPU and NUMA placement to follow the PF. Reviewed by: jhb MFC after: 2 weeks Sponsored by: BBOX.io Differential Revision: https://reviews.freebsd.org/D59062
w83793g_writereg is unused so silence it. Fixes: https://cgit.freebsd.org/src/commit/?id=cd3cc6e910c0f ("i2c/sensors: Add driver for W83793 hardware monitor")
This addresses a bug in the addition overflow check added in commit
319414a926af ("netmap: Handle overflow when computing ring sizes"): that
overflow wasn't actually caught by the check because "len" is promoted
to size_t.
PR: https://bugs.freebsd.org/bugzilla/show_bug.cgi?id=297300
Fixes: https://cgit.freebsd.org/src/commit/?id=319414a926af ("netmap: Handle overflow when computing ring sizes")
Sponsored by: The FreeBSD Foundation
Differential Revision: https://reviews.freebsd.org/D58896
acpi_pci: Preserve CPU locality queries for descendants bus_generic_get_cpus() preserves the original leaf device while forwarding a request through the bus hierarchy. Consequently, acpi_pci_get_cpus() may receive a descendant below a PCI function rather than one of the PCI bus's direct children. Only apply the SR-IOV PF-locality mapping to direct PCI children. Preserve the previous ACPI CPU-locality lookup for descendants so their unrelated bus ivars are not interpreted as PCI device information. Reviewed by: jhb MFC after: 2 weeks Sponsored by: BBOX.io Differential Revision: https://reviews.freebsd.org/D59206
acpi_pci: Cache PCI proximity domains A PCI function's _PXM is stable for the lifetime of its device instance, but CPU and DMA locality queries may evaluate it repeatedly. SR-IOV amplifies this because every VF resolves locality through the same PF. Cache successful mappings and the stable absence of _PXM on the locality source device, and share that result between CPU and domain queries. Continue to retry generic evaluation or mapping errors rather than making a potentially transient failure permanent. Reviewed by: jhb MFC after: 2 weeks Sponsored by: BBOX.io Differential Revision: https://reviews.freebsd.org/D59207
acpi_pci: Honor device proximity for DMA tags A PCI function with its own _PXM still inherits a DMA tag carrying the upstream bridge's proximity domain. Resolving an SR-IOV VF's locality through its PF therefore does not affect the domain used for DMA allocations. Create and cache a private child tag when the function, or a VF's owning PF, has an explicit _PXM. Parent it to the existing PCI or IOMMU tag so its constraints remain intact, then apply the function's domain without mutating a shared tag. pci_get_dma_tag() already performs the IOMMU lookup, so remove the duplicated lookup in the ACPI subclass while here. Reviewed by: jhb MFC after: 2 weeks Sponsored by: BBOX.io Differential Revision: https://reviews.freebsd.org/D59063
rangelock: Fix format strings for 32-bit kernels Reported by: Jenkins Fixes: https://cgit.freebsd.org/src/commit/?id=f1f58bdf7b5f ("acpi_pci: Honor device proximity for DMA tags")
These helper function will not be needed anymore because the way that clocks are parsed from the DTS is reworked. Approved by: imp, manu(mentor) Tested by: Rick Richard Differential revision: https://reviews.freebsd.org/D46712
Due to upstream has changed to use clock-output-names we need to update the names in our code aswell. Approved by: imp(earlier revison), manu(mentor) Tested by: Rick Richard Differential revision: https://reviews.freebsd.org/D46713
The Atom C2000 integrated GbE programming reference documents the I354 internal memory error architecture. It shares the I350 PEIND and ICR.FER routing, DMA and packet-buffer status, LAN parity status, and required reset recovery. Extend the existing I350 recovery and corrected error accounting paths to I354. Keep the PCIe corrected error mask family-specific. C2000 PCIEECCSTS ends at the transmit write-data indication in bit 4 and does not implement the I350 retry buffer indication in bit 5. Do not expose the corresponding retry counter on I354. The PRM overview says a PCIe region failure requires a system reboot, while the individual PCIEERRSTS fields prescribe CTRL.RST followed by port reinitialization. Use the register specific recovery, matching the existing I350 path; failed reinitialization still leaves the port down. This follows sections 5.6 and 6.21 of the Intel Atom Processor C2000 Product Family Integrated GbE Controller Programmer's Reference Manual, document 537426 revision 1.5. MFC after: 2 weeks Sponsored by: BBOX.io
I350 and I354 report a corrected ECC error in the LAN transmit management FIFO through LANPERRSTS bit 16. Unlike the parity status in the same register, this condition neither interrupts nor stops traffic. Poll the latch with the other corrected error status, increment a dedicated counter, and clear only its RW1C bit. Expose it as dev.igb.N.memory_errors.corrected_lan_mng_fifo. Fatal error handling returns before the periodic statistics sweep and may reset the device. Drain all I350 and I354 corrected-error status in the admin task before recovery so the reset does not discard pending indications. This follows section 6.21.16 of the Intel Atom Processor C2000 Product Family Integrated GbE Controller Programmer's Reference Manual, document 537426 revision 1.5. MFC after: 2 weeks Sponsored by: BBOX.io
To date, pmap_update_entry() has unconditionally passed false as final_only to pmap_s1_invalidate_range(). Passing false means that we invalidate the intermediate "page walk cache" entries in the TLB as well as the leaf that is being replaced. However, invalidating intermediate entries is only necessary when doing a superpage promotion that replaces a pointer to a page table page by a large page mapping. Reviewed by: andrew, kib, markj MFC after: 1 month Differential Revision: https://reviews.freebsd.org/D58917
The disabled path wrote the complement of DMAC_EN to DMACR. That set every other field, including reserved bits, the one-shot EXIT_DC command, watchdog enables, receive threshold, and PCIe Lx selection. When DMA coalescing was enabled, the requested watchdog and Lx-delay values were ORed into their reset values rather than replacing the fields. PCIEMISC.LX_DECISION was also cleared, preventing DMACR.DMAC_Lx from controlling PCIe low-power entry. Disable coalescing by clearing only DMAC_EN and retaining the documented DMAC_Lx policy and watchdog fields. On enable, replace the variable fields under their masks, select DMA requirements for PCIe low-power entry, and program the loopback and BMC watchdog policies independently of prior state. Make igb_init_dmac() the sole owner of this policy, including when SR-IOV is active, so later IOV initialization cannot overwrite it. Apply this consistently to I350, I354, and I210 while preserving the I354-specific timer units. Leave I210 reserved fields at their required encodings and do not expose the ineffective control on I211. Validated on I210 and I350 hardware. Repeated I210 enable, disable, and reset cycles preserved the watchdog, Lx, TTLX, and LX_DECISION fields. On I350, dmac values of 250, 1000, and 10000 programmed DMACWT as 0x7, 0x1f, and 0x138, respectively, while retaining a four-tick TTLX and leaving the reserved and watchdog fields stable. A real reset initiated by a sibling function restored the enabled dmac=1000 tuple and a down interface path restored the disabled tuple. MFC after: 2 weeks Sponsored by: BBOX.io
led(4) invoked the led_t callback with its mutex held, including from the blink callout, so the callback must not sleep. PR: https://bugs.freebsd.org/bugzilla/show_bug.cgi?id=251032 MFC After: 1 week Reviewed by: adrian Differential Revision: https://reviews.freebsd.org/D59263
CTRL.DEV_RST resets every port on an 82580 and newer igb device. Hardware reports the event to each affected function through ICR.DRSTA and requires software to reinitialize the port registers and descriptor rings. The driver neither enabled nor handled this cause, so a reset initiated by another function could leave a running interface using stale state. In FreeBSD, we do not currently send this, but other OSes including Linux do, so a PF passed through to such a guest or FreeBSD as a guest running with a passthrough PF on the same controller can be wedged. Enable DRSTA for 82580 and newer PFs in both MSI-X and shared MSI/legacy modes. Bit 30 is reserved on 82575 and is the TCP timer on 82576, so leave it masked on those parts. Latch the event without programming port registers from the interrupt filter. Defer the iflib reset request to admin-task context because the request takes STATE_LOCK. Use IAM to auto-mask shared interrupts on the first ICR read. Keep admin, queue-vector, and interrupt-rearm work quiesced until initialization succeeds. Follow the device reset handshake before the first CTX-owned register programming: wait for GCR to report that the device reset and pending PCIe transactions have completed, then acknowledge STATUS.DEV_RST_SET. Also verify EEPROM autoload and PF reset completion on I350 and newer parts. This follows the Intel 82580 Datasheet, section 4.3.2, and the Intel I350 Datasheet, section 4.3.4. Detect a completed reset from STATUS even when the interface was down. A bounded timeout falls back to the requested port reset. Retain staged reset state through the complete initialization. Check ICR.DRSTA, GCR, and STATUS.DEV_RST_SET after all registers and rings have been rebuilt. If another reset arrived while interrupts were masked, reject the incomplete initialization and repeat the handshake. If both MMIO and PCI configuration space have disappeared, leave the interface stopped rather than queueing an endless recovery loop. Validated on an I210 with INVARIANTS and WITNESS. Injecting CTRL.DEV_RST while igb0 was running recovered through a full iflib initialization in both MSI-X and MSI modes without a panic. STATUS.DEV_RST_SET and GCR.DEV_RST_IN_PROGRESS cleared, enabled DMA coalescing was restored. Injecting the reset while igb0 was down left DEV_RST_SET latched; the first ifconfig up consumed it. On a dual-port 82580, synthetic DRSTA injection produced one complete reinitialization in four-queue MSI-X and shared-MSI modes. Five repeated events produced five clean reinitializations without a panic or watchdog. A raw CTRL.DEV_RST test is not counted because Intel's shared code deliberately avoids that unreliable operation on 82580. On an I350, a real device-wide reset initiated by a sibling function produced one reinitialization while igb0 was running and restored its carrier and enabled state. With igb0 down, STATUS.DEV_RST_SET remained latched until the first up, which consumed it and restored the correct state. The host remained healthy in both cases. MFC after: 2 weeks Sponsored by: BBOX.io
Get rid of the powernv-specific AP callback and use the new-ish PIC_AP_INIT() PIC KPI instead.
The loop goes over the qman channel total (16), so if a port ID is not found in the list it could walk off the end of the list and return garbage. Not a problem in practice, as only valid ports are included in our device trees, but protect it anyway.
Instead of forcing an `mdio` pseudo-device to hang off the xmdio, rename xmdio to "mdio" and make it an ofw bus device, akin to the mii_fdt driver, so that children can get the device tree goodies.
The presence of the PCI power management capability does not imply that a function can signal PME# from every power state. Drivers which advertise wake based only on pci_has_pm() can consequently expose wake modes that cannot work. Add pci_has_pme() to query the PME_Support bitmap for a specific state. Use it to implement LinuxKPI pci_pme_capable(), removing its duplicate PME_Support decoder. Validated the helper against PCI PMC capability values from 82571EB, 82573L, 82579LM, I210, I225, and I226-V controllers. The 82571 and 82573 reported PMC 0xc822, while the I226-V reported 0xc823. In both values, bits 15, 14, and 11 advertise PME from D3cold, D3hot, and D0; the low-bit difference is only the PM capability version. MFC after: 2 weeks Sponsored by: BBOX.io
The attach path translated WUC.APME into a saved link-change filter, then advertised magic-packet wake. Suspend removed unselected magic, unicast, and multicast bits from that saved value, commonly leaving no hardware wake filter at all. The destructive masking also made later capability changes ineffective. Advertise the I225/I226 wake filters whenever PCI power management is available and enable magic-packet wake by default. Build a fresh WUFC mask for every suspend, and explicitly clear WUC, WUFC, and PCI PME when wake is disabled. Require the PCI power-management capability to report D3hot PME support before advertising or arming wake. A PM capability alone does not mean the function can signal PME from the state used during system sleep. Reconstruct RAR0, the multicast table, and the receive filter after the stop-time reset so unicast and multicast wake use the current interface state. Keep the PHY powered while wake is armed, drain pending PCIe transactions, and disable bus mastering before D3. Clear the sticky base and extended wake status before arming filters. On resume, report the saved hardware wake cause, clear the device wake source, and then clear PCI PME. Run the DMA-fencing sequence even when wake programming fails, while allowing shutdown to continue after logging the failure. Explicitly restore PCI bus mastering during initialization so an iflib-local resume after another child rejects suspend can restart the device. Restore the Intel shared code PHY power down test from DPDK. The FreeBSD split inverted the reset-block condition, so its dormant hook would power down only when firmware explicitly vetoed the operation. Do not copy the generic legacy management test verbatim. The I225 and I226 define MANC bits 0 and 1 as flow-control and NC-SI discard controls, not SMBus and ASF enable bits. Preserve the link when the documented TCO receive path is enabled; the shared reset-block test separately honors the firmware keep-link-up veto. This follows section 8.21.1 of the Foxville Software User Manual. Evaluate management pass-through at each suspend. When neither a host wake filter nor management requires the link, invoke the shared-code power-down hook before D3. Leave a management-owned link untouched and restore a link previously powered down by the driver without another PHY reset. This follows DPDK's stop/start pairing without changing ordinary ifconfig down behavior. The implementation was checked against the Intel I225/I226 programming model, the Intel Linux igc lifecycle, and DPDK. On an I225-LM, a device-only D3 test observed PME for a valid magic packet and no PME with every host wake filter disabled. Both cases resumed to D0 with the link and configured addresses operational. The I225 system also completed a full ACPI S3 cycle and resumed with link, configured addresses, and traffic operational. On an I226-V, a device-only D3 test changed PMCSR accordingly after a magic packet. With dev.igc.0.wake enabled, a full ACPI S3 cycle remained asleep until a delayed magic packet and resumed with link, addresses, and traffic operational. PR: https://bugs.freebsd.org/bugzilla/show_bug.cgi?id=282140 Obtained from: DPDK (shared-code power-down structure) MFC after: 2 weeks
The driver used the NVM APME default as both the hardware-support decision and the mutable filter mask. Consequently, an NVM-disabled but capable port did not advertise wake support, disabling a wake mode once could keep it disabled across later suspends, and directed-unicast wake could never be selected. Require the PCI power management capability to report D3hot PME support before advertising or arming wake. A PM capability alone does not mean the function can signal PME from the state used during system sleep. Separate the board and port capability matrix from the NVM-selected magic packet default. Read the proper per function NVM word on igb controllers, cover the newer PCH generations, and retain the documented legacy, multi-port, and OEM restrictions. Decode the distinct APM Enable locations used by 82544, 82541EI/82547EI, and the later 8254x parts. Do not advertise wake on the 82541ER, whose power-management logic cannot assert PME for wake events. For I210/I211 internal iNVM, use the hardware-loaded WUC.APME state; the shared reader does not expose the optional Initialization Control 3 word. Build WUFC from the enabled ifnet capabilities for each suspend. Reconstruct RAR0, the multicast table, and the receive filter after the stop-time reset so unicast and multicast wake use the current interface state. Fill the MTA on legacy PCI/PCI-X controllers and 82575 through 82580 when the address list overflows; their multicast wake matchers require the indexed MTA bit and do not use RCTL.MPE as a substitute. Do not access PF-only wake CSRs from the igb VF suspend and resume paths. Use the shared BM page access helpers and propagate every PHY receive address and wake-register programming failure. On resume, perform the required LCD reset before clearing host PHY-wake ownership, report the saved PHY or MAC wake cause, and clear PCI PME after removing the device wake source. Preserve management engine wake ownership throughout. Keep WUC.APME set only when early 82545EM/82546EB manageability needs its D3 clock-tree workaround; ordinary host wake uses PCI PME. Keep the link powered while host wake is armed. With no host wake, evaluate management pass-through at each suspend. Leave a management-owned link untouched and keep PCI PME enabled. Otherwise, use the Intel shared code PHY power-down hook, or its matching SerDes shutdown hook on igb fiber and SerDes devices. Track that state and restore the link without another PHY reset before hardware initialization. Ordinary ifconfig down behavior is unchanged. Apply and undo the PCH Sx workarounds across their full supported range. Use controller-specific CTRL and laser semantics, and restore RCTL when wake setup fails. Always run the pending PCIe-transaction drain and bus-master-disable sequence before D3. Suspend reports a wake programming failure rather than sleeping without wake, shutdown logs it and continues through the fencing sequence. Do not apply the ICH/PCH IGP3 D3 power-down workaround to igb controllers. The merged driver inherited an unconditional call from the em-only driver. On 82575 and 82576 it asserted CTRL.PHY_RST after the wake filters were armed, preventing the link from receiving wake traffic. The implementation was checked against the Intel controller data sheets, the Intel Linux e1000, e1000e, and igb lifecycle code, DPDK, and the Intel FreeBSD em-7.7.8 and igb-2.5.31 drivers. The 8254x audit also covered the PCI/PCI-X Software Developer's Manual, the 82541/82547 NVM guide, and the 82544, 82545, and 82546 specification updates. The out of tree drivers carry the family-specific power down and reset block hooks but do not call them from suspend. DPDK supplies the stop/start pairing. On PCH controllers including an 82579LM, I217-LM, and various I219s, device-only D3 tests observed PME and BM_WUS.MAG for a magic packet, no PME with every host filter disabled, and BM_WUS.EX with only directed-unicast wake enabled. With dev.em.0.wake enabled, ACPI S3 slept until a delayed magic packet and resumed with the interface operational. After wake traffic stopped and resume completed, a second cycle again waited for a newly delayed magic packet. Traffic restored after both host and firmware wake were enabled. An 82574L woke from S3 after one magic packet, reported MAC wakeup status, and returned with link and traffic operational. On 82571EB and 82573L adapters, device-only D3 tests observed WUS.MAG and PCI PME status after a magic packet, then returned to D0 with link and traffic operational. Full S3 did not wake either add-in card. The positive device tests and negative S3 isolate the remaining failure outside the MAC filter programming and my cards may lack aux power wiring because the link was off in S3. On 82575EB and 82576 adapters, pre-fix device only D3 tests left PMCSR at 0x2103 despite ten verified magic packets, and the handoff showed CTRL.PHY_RST asserted. With the em-family gate, identical tests changed PMCSR from 0x2103 to 0xa103, resume reported WUS.MAG, and both links returned operational. S3 testing on these separated controller from board behavior. An Intel 82576 card retained link in S3 and woke the system from a delayed magic packet, reported WUS.MAG, and returned with interface operational. The tested 82575 add-in card lost its link LED in S3 and retained no WUS cause after manual resume, although its identical D3hot test passed. That points the 82575 S3 result to card aux power wiring as well. D3 tests were performed on I210 and I350 but S3 has not yet been attempted on them. lem(4) testing has not been attempted yet. Community reports of success and failure are welcome. PR: https://bugs.freebsd.org/bugzilla/show_bug.cgi?id=232708, https://bugs.freebsd.org/bugzilla/show_bug.cgi?id=238411, https://bugs.freebsd.org/bugzilla/show_bug.cgi?id=295443, https://bugs.freebsd.org/bugzilla/show_bug.cgi?id=296675 MFC after: 2 weeks Sponsored by: BBOX.io
The PCH suspend path kept a wake link fully powered and did not restore the negotiated EEE modes after its stop-time reset. Intel provides the ULP entry and exit machinery in the shared code, but FreeBSD did not invoke its Sx policy. Enter ULP on LPT and newer PCH controllers when wake is armed without directed-unicast, multicast, or broadcast filters, which ULP cannot preserve. For a link retained by host wake or management, restore the 100BASE-TX and 1000BASE-T LPI controls selected by the local advertisement and the cached link-partner ability. Keep these power reductions best-effort: wake filters and PME are already configured independently, and a ULP or EEE failure is logged without converting an optional power optimization into a suspend failure. The existing PCH resume workaround forcibly exits ULP and clears automatic Sx LPI state before normal initialization. Validated on a ThinkPad T440p with an I217-LM. FBT confirmed that the helper received the magic-wake mask during device suspend. The shared ULP helper returned its documented no-op for the initial I217 device ID. A full S3 cycle resumed cleanly. On a ThinkPad P51 with an I219, full S3 waited for a magic packet and then resumed with the 1-Gbps link and traffic restored. MFC after: 2 weeks Sponsored by: BBOX.io
The 82576 and I350 retain VF queue enable and DMA address state across VFLR. iflib enables PCI bus mastering before driver attach, so stale state left by a previous owner can otherwise issue DMA before igbvf has completed its first reset and queue sanitization. Disable PCI bus mastering immediately after mapping the VF BAR. Keep it disabled until reset and queue sanitization succeed, verify both disable and enable through PCI command-register readback, and wait for pending transactions before treating the fence as complete. Resanitize on stop before iflib releases queue mappings. The sanitizer and recovery were exercised on I350 and 82576 VFs. Forced queue-disable failure left the VF down, and a later administrative down/up recovered it; successful I350 VFs passed bidirectional traffic with no errors or drops. Sponsored by: BBOX.io
FreeBSD's bxe hardwires CNIC_SUPPORT() to 0, so bxe_ilt_set_info() never enters the block that initializes the SRC and TM ILT clients. Those two clients are left zeroed (page_size 0, flags 0), yet ecore_ilt_init_page_size() calls ecore_ilt_init_client_psz() for all four clients unconditionally. For SRC and TM that evaluates ILOG2(page_size >> 12), i.e. ilog2(0). On an INVARIANTS kernel ilog2() asserts "ilog argument must be nonzero" and panics the machine the first time the interface is brought up (bxe_init -> bxe_nic_load -> bxe_init_hw -> ecore_ilt_init_page_size). On a non-INVARIANTS kernel it silently programs a bogus page-size register instead. Restore the else branch that upstream Linux bnx2x carries: when CNIC is not supported, mark the SRC and TM clients with ILT_CLIENT_SKIP_INIT and ILT_CLIENT_SKIP_MEM so ecore_ilt_init_client_psz() skips them. Root-caused from a crash dump on a BCM57810 (device 0x168e): the ILT clients showed CDU and QM populated and SRC and TM zeroed with no skip flag set. Reviewed by: adrian Approved by: adrian (mentor) Differential Revision: https://reviews.freebsd.org/D58587 Signed-off-by: Nick Price <nprice@FreeBSD.org>
clk_cpll_div_333m_div, clk_cpll_div_125m_div, clk_cpll_div_50m_div, clk_cpll_div_25m_div, clk_cpll_div_100m_div, clk_osc0_div_750k_div did not respond Rockchip RK3568 TRM Part1 V1.1-20210301.pdf documentation page 79. I changed them correctly. Reviewed by: imp Pull Request: https://github.com/freebsd/freebsd-src/pull/2287
tpm: Correct the TPM 1.2 suspend transaction The legacy driver wrote TPM_ORD_SaveState directly to the command FIFO, but used ordinal 156 instead of the TPM 1.2 ordinal 152 and never completed the transaction through the transport start and end methods. On a TIS device this omitted TPM_STS_GO, and the response read used the header length as flags instead of requesting the complete parameter size. The legacy Atmel reader would also dereference the null byte-count pointer. Send the header-only command through the normal transport lifecycle, validate the response header and TPM result, and retry TPM_WARN_RETRY for a bounded five seconds. Fail suspend rather than enter S3 after an unsuccessful state save. This follows the TPM 1.2 SaveState command definition and the bounded retry policy used by other TPM 1.2 implementations. The stock driver failed to resume a ThinkPad T440p with its STMicro TPM 1.2 Security Chip enabled; disabling the chip made S3 reliable. With this change and the following TIS resume restoration, the enabled TPM completed two consecutive S3 cycles. PCR 0 was readable with the same value before and after each cycle, and no TPM errors were logged. PR: https://bugs.freebsd.org/bugzilla/show_bug.cgi?id=291067 Reviewed by: kevans MFC after: 2 weeks Sponsored by: BBOX.io Differential Revision: https://reviews.freebsd.org/D59192
tpm: Restore TPM 1.2 TIS state after resume Firmware restores the state saved by TPM_ORD_SaveState, but the TIS interrupt, locality, and command FIFO state are not guaranteed to survive S3. The legacy driver previously treated resume as a no-op. Revalidate the interface and device identity, disable and acknowledge stale interrupts, restore the configured interrupt vector, reacquire locality zero, and return the FIFO to command-ready state. Also disable TIS interrupts during initial setup when the device uses polling so firmware settings cannot leave an unhandled interrupt enabled. TIS 1.3 Table 22 makes the interrupt control registers locality protected. Acquire locality before disabling or programming them during initial setup and resume rather than relying on probe retaining locality. Keep TPM self-test outside the resume critical path. It can take minutes on some TPM 1.2 devices and is not required to restore the transport state. The two-commit suspend and resume series completed two consecutive S3 cycles on a ThinkPad T440p with its STMicro TPM 1.2 Security Chip enabled. PCR 0 was readable with the same value before and after each cycle, and no SaveState or TIS restoration errors were logged. The locality ordering completed another two consecutive S3 cycles on a ThinkPad T430 with the same STMicro TPM in polling mode. PCR 0 again remained stable, and TPM access recovered without errors after each resume. PR: https://bugs.freebsd.org/bugzilla/show_bug.cgi?id=291067 Reviewed by: kevans Sponsored by: BBOX.io MFC after: 2 weeks Differential Revision: https://reviews.freebsd.org/D59193
tpm: Remove Giant from the TPM 1.2 driver Serialize TPM 1.2 commands, character-device methods, and power transitions with an sx lock, following the command ownership model used by the TPM 2.0 driver. Reject new operations once detach starts and drain the character device before releasing transport resources. Giant also closed the interrupt race between the final TIS status check and tsleep. Replace that implicit dependency with a mutex and condition variable, use an absolute deadline across unrelated wakeups, and make the interrupt handler MPSAFE. Create the device node atomically with its softc and finish failed write transactions so every command path releases its transport state. The polling path was validated on ThinkPad T430 and T440p systems with their STMicro TPM 1.2 devices enabled. Exclusive-open behavior, 100 consecutive PCR reads, and module unload and reload completed without errors on both systems. Two consecutive S3 cycles on each system preserved PCR values and command access, including another 100 PCR reads after resume, without lock or TPM diagnostics. Reviewed by: kevans, seuros MFC after: 2 weeks Sponsored by: BBOX.io Differential Revision: https://reviews.freebsd.org/D59211
tpm: Bound TPM 1.2 locality ownership A TIS locality must remain active while a command is in flight, but should be relinquished once the command completes or is abandoned. The driver retained locality zero after probe, initialization, and resume, and several transaction error paths returned without releasing it. Closing the device after writing a command without reading its response had the same effect. Track locality ownership and whether a command is awaiting its response. Release locality after probe, initialization, and resume; retain it only across a successful command write and its matching response read. Abort and release on errors, replacement commands, close, and detach. Wait for locality during ISA probe instead of assuming an immediate grant, release locality acquired by the probe, and stop treating the command-style TPM_ACCESS register as restorable state. On a ThinkPad T440p with an STMicro TPM 1.2, the old driver left TPM_ACCESS at 0xa1 immediately after attach. The new driver left it at 0x81 after attach, completed PCR reads, and closing with an unread response. PCR reads also survived an unload and reload without TPM or locking diagnostics. Reviewed by: kevants MFC after: 2 weeks Sponsored by: BBOX.io Differential Revision: https://reviews.freebsd.org/D59237
tpm: Move user copies outside the TPM 1.2 lock The character-device paths held the transaction and lifecycle lock while uiomove() accessed user memory. A user page fault could therefore delay suspend or detach, and a copyout failure occurred while the TPM response was still active. Copy commands into the bounded stack buffer before taking the lock. For reads, validate the response header, buffer the complete response while the lock is held, finish the TPM transaction, and copy it to userspace after unlocking. Use a non-blocking allocation so memory pressure cannot turn response buffering into another lifecycle wait. NetBSD uses the same separation but limits responses to its fixed 1 KiB buffer. Allocate the TPM-advertised response length to preserve the existing FreeBSD support for larger streamed responses. On a ThinkPad T440p with an STMicro TPM 1.2, a PCR read into a 4 KiB userspace buffer returned the expected 30-byte response. A deliberately short five-byte read failed cleanly, relinquished locality zero, and the next PCR read succeeded. Reviewed by: kevans MFC after: 2 weeks Sponsored by: BBOX.io Differential Revision: https://reviews.freebsd.org/D59238
mmu_radix_sync_icache() walked the page tables with an unlocked
pmap_extract() and passed the result straight to PHYS_TO_DMAP(),
checking only that it was non-zero. Nothing keeps the mapping - or the
page table page holding it - alive across that window: if another thread
of the same process tears a mapping down concurrently, the page table
page can be freed and reused, so pmap_extract() reads arbitrary memory
and returns a bogus physical address. __syncicache() then dereferences
an unmapped direct map address and the kernel takes a data storage
interrupt:
fatal kernel trap:
exception = 0x300 (data storage interrupt)
virtual address = 0xc003317ca6022a00
dsisr = 0x40000000
srr0 = 0xc000000000f59460 (__syncicache)
lr = 0xc000000000f23588 (mmu_radix_sync_icache)
pid = 23878, comm = skyframe-evaluator-
panic: data storage interrupt trap
The faulting addresses decode to physical addresses far beyond installed
memory (~140 TB and ~900 TB on a 256 GB machine), i.e. translations that
never existed.
The hash MMU implementation of the same method, moea64_sync_icache(),
already holds PMAP_LOCK() across the loop; do the same here.
mmu_radix_extract() does not acquire the pmap lock itself, so this
introduces no recursion.
JIT workloads reach this path constantly: ppc_instr_emulate() calls
pmap_sync_icache() on the faulting address for the SIGILL "second
chance" retry, so a multithreaded JVM executing freshly written code
races against its own threads' mmap/munmap. Every panic observed here
was in a JVM thread.
Tested on POWER9 (radix MMU) with a bazel/JVM build loop that previously
panicked the machine twice within ten minutes: afterwards 13 consecutive
builds and more than 10 hours of uptime with no panic, on both
15.1-RELEASE and 16.0-CURRENT.
MFC after: 1 week
Differential Revision: https://reviews.freebsd.org/D59311
Reviewed by: jhibbits, adrian
As noted in the comment, some headsets with a hardware sidetone are incredibly sensitive and emit immediate feedback upon attach with the current system-wide default of 75%. Drop it down just for snd_uaudio(4) to avoid incredibly unpleasant surprises. MFC after: 3 days Reviewed by: christos Differential Revision: https://reviews.freebsd.org/D59199
pmc_save_user_callchain() emits the pc it just loaded before checking whether fp is the ABI's zero frame-chain terminator. At the bottom of a well-formed chain under _start, fp comes back 0 as expected, but the paired pc is stale rtld data left on the stack -- a legal userspace VA that still passes PMC_IN_USERSPACE(), so it gets emitted as a bogus extra frame. This shows up in flame graphs as a spurious hex-valued root frame below _start. Check fp == 0 alongside the existing checks before emitting, matching how arm/arm64/powerpc already load the next fp before their check. Measured via 1kHz hwpmc sampling on an OCA: stacks with any unresolved hex frame drop from 23.9% to 1.3%, and stacks with hex at the root drop from 5.5% to 0.3%. Reviewed by: mhorne, Ali Mashtizadeh <ali@mashtizadeh.com>, gallatin MFC after: 3 days Sponsored by: Netflix Differential Revision: https://reviews.freebsd.org/D59229
Reviewed by: fuz Approved by: fuz (mentor) MFC after: 1 month Differential Revision: https://reviews.freebsd.org/D59293
Mark the driver stopped after attach so its first IFDI_STOP() call does not repeat hardware shutdown. Defer error interrupt recovery through iflib instead of calling driver stop and init methods from interrupt context. Let iflib own the stop and restart around MTU changes as well, avoiding duplicate lifecycle operations. MFC after: 2 weeks Sponsored by: BBOX.io
Register the standard iflib device methods for shutdown, suspend, and resume. This gives axgbe the framework managed reinitialization used by other iflib drivers after a power transition. MFC after: 2 weeks Sponsored by: BBOX.io
Register the iflib device suspend and resume methods so the existing driver callbacks run during system power transitions. This stops mailbox retry work before suspend and lets iflib reinitialize the datapath after resume. MFC after: 2 weeks Sponsored by: BBOX.io
Register the standard iflib device suspend and resume methods so the framework reinitializes the VF datapath after a system power transition. MFC after: 2 weeks Sponsored by: BBOX.io
Register the iflib device suspend and resume methods. Remove the direct initialization from the driver resume callback because iflib_device_resume() performs the datapath restart after the callback returns. MFC after: 2 weeks Sponsored by: BBOX.io
The internal TPM2_Shutdown and TPM2_Startup paths ignored both transport failures and the TPM response. Suspend could therefore enter S3 without saved TPM state, while resume could restart entropy harvesting after a failed state restoration. Build both commands through one helper, validate their response framing and TPM return codes, and propagate failures. Retry the standard RETRY and TESTING responses with bounded exponential backoff. Accept TPM_RC_INITIALIZE from Startup because firmware may already have started the TPM during resume. Do not enter S3 after an unsuccessful state save, and do not restart the entropy task when TPM state restoration failed. If Shutdown fails after the entropy task was drained, requeue it before returning so an aborted suspend does not permanently stop harvesting. Reviewed by: kevans MFC after: 2 weeks Sponsored by: BBOX.io Differential Revision: https://reviews.freebsd.org/D59195
The TIS attach path tested its interrupt by transmitting GetRandom before tpm20_init() allocated the internal command buffer. A TPM2 FIFO device with a usable IRQ could therefore dereference a null internal_priv. Initialize the common TPM2 state before running the interrupt test. Make common cleanup safe for partially initialized devices and leave cleanup to the attachment after tpm20_init() fails, avoiding duplicate release of the lock, command buffer, and random-source state. Clear the IRQ resource pointer after releasing it when interrupt handler setup fails so the later polling-mode detach does not release it twice. Free the internal command allocation through its object pointer rather than relying on its embedded buffer being the first structure member. Reviewed by: kevans MFC after: 2 weeks Sponsored by: BBOX.io Differential Revision: https://reviews.freebsd.org/D59196
TIS interrupt routing and enable registers may lose their state across S3, while the driver retains its software indication that interrupts work. A subsequent locality or command wait can then sleep for an interrupt that cannot arrive. Remember whether interrupts worked before suspend and restore the vector, pending status, and enable mask before TPM2_Startup. Put the transport in polling mode first; the interrupt handler promotes it back to interrupt waits only after observing an interrupt from the restored configuration. If register restoration fails, Startup and subsequent commands continue using polling. Preserve the initial interrupt-enable mask, including the firmware's trigger and polarity selection proven by the attach time interrupt test, and restore that exact mask rather than accepting post-S3 defaults. Program the same safe baseline for polling devices during attach and resume. Acquire locality, disable global interrupt delivery, and acknowledge pending status so firmware cannot leave interrupts armed without a handler. Use the same register programming helper during attach and resume, and stop trying to configure interrupts after a locality acquisition failure. Three consecutive device suspend and resume cycles completed on a Lenovo TPM2 FIFO device without an IRQ resource. GetRandom succeeded after each resume, and module detach completed without errors. Reviewed by: kevans MFC after: 2 weeks Sponsored by: BBOX.io Differential Revision: https://reviews.freebsd.org/D59197
Mark the device as dying before teardown and destroy the character device before freeing its private state or lock. This prevents cdev methods from entering with a freed internal buffer or a destroyed sx. Check the teardown state in command paths, honor failures from the cdev private data interface, and publish teardown before waiting for the lifecycle lock. Keep that lock across TPM retry delays so commands cannot interpose and private state remains pinned, but abort before the next retry once teardown begins. Block new cdev operations after a successful Shutdown(STATE). Keep the suspend gate and the TPM command under the same lock so a userspace command cannot invalidate the saved state before S3 entry. Clear the gate only after Startup(STATE) succeeds. Keep entropy harvesting scheduled after a transient command or suspend failure, but stop it while suspended or once teardown begins. Queue the next timeout while holding the lifecycle lock so release cannot miss a concurrent requeue. Validated on two TPM 2.0 FIFO systems. Each completed five device suspend/resume cycles, rejected both new and already-open cdev operations with EBUSY while suspended, completed 200 concurrent PCR reads, and detached cleanly while four PCR readers were active. A ThinkPad P51 also completed a full S3 cycle with PCR 0 unchanged and 50 successful reads after resume. Reviewed by: kevans MFC after: 2 weeks Sponsored by: BBOX.io Differential Revision: https://reviews.freebsd.org/D59240
The TIS interrupt handler can acknowledge and signal an event after the waiter checks the device status but before it enters tsleep(). Since the handler is MPSAFE, the command lock does not close this window. A lost wakeup can delay a completed command for its full timeout, up to 40 seconds for long TPM 2.0 operations. Publish the expected event under an interrupt mutex and use a generation counter to record matching interrupts. Recheck the device predicate without the mutex because register access may sleep on a SPI transport, then compare the generation before atomically waiting on a condition variable. This closes the check-to-sleep race without placing sleeping bus operations under a mutex. Use an absolute deadline while retrying the predicate after wakeups. Apply the same scheme to locality acquisition, which had an equivalent race. Leave the expected event published while polling so the attach-time test can still prove that an advertised interrupt arrived. Regression-tested the polling fallback on two TPM 2.0 FIFO systems with 200 concurrent PCR reads per system and repeated device suspend/resume. Neither ACPI device exposes an IRQ, so the interrupt-mode path remains hardware unvalidated. Reviewed by: kevans MFC after: 2 weeks Sponsored by: BBOX.io Differential Revision: https://reviews.freebsd.org/D59241
tpm20: Release transport state after command failures Once a transport acquires locality, several TIS and CRB error paths return without relinquishing it. They can also leave a partial FIFO transaction or an active CRB command for the next operation to inherit. Route post-locality exits through common cleanup. Reset the TIS command state on every attempt. For CRB, cancel an active failed command when necessary, request the idle state, and relinquish locality even when the state transition itself fails. Successful command handling is unchanged apart from sharing the same cleanup path. Reviewed by: kevans MFC after: 2 weeks Sponsored by: BBOX.io Differential Revision: https://reviews.freebsd.org/D59243
tpm20: Correct 32-bit register helpers OR4() reads only the low byte before writing the complete 32-bit register. Preserve all register bits by using a matching 32-bit read. Make BIT() produce an unsigned value so masks containing bit 31 do not rely on a signed left shift into the sign bit. OpenBSD carries the same change. Reviewed by: kevans MFC after: 2 weeks Sponsored by: BBOX.io Differential Revision: https://reviews.freebsd.org/D59244
tpm20: Move user copies outside the lifecycle lock The TPM 2.0 character-device methods held the global device lock while uiomove() accessed user memory. User page faults could therefore delay suspend or detach even though the read response was already buffered. Add a per-open sleepable lock to serialize operations on each response buffer. Stage commands under that lock before acquiring the device lock, and copy them into the response buffer only after the lifecycle checks succeed. This preserves an unread response when suspend or detach rejects a write. Release the device lock before copying buffered responses out. Also advance the response offset by the bytes actually copied when uiomove() returns after a partial transfer. Validated on an Intel TPM 2.0 TIS device. PCR reads and GetRandom passed under 16-process mixed command load. A response was consumed correctly in 5-byte, 7-byte, and remainder reads. Module unload/reload recreated the device and entropy source without lock diagnostics. Source inspection confirmed rejected writes preserve unread responses. Reviewed by: kevans MFC after: 2 weeks Sponsored by: BBOX.io Differential Revision: https://reviews.freebsd.org/D59245
Move the dtrace_trap hook at the start of the abort handler to exit early when a trap is handled by DTrace. Fix the type argument to be the actual fault type instead of the value of the FAR. The latter will need to be added to the trapframe, until then DTrace will report unmapped addresses as the null address. Correct the comment of the PUSHFRAMEINSVC assembler macro to reflect that coming from SVC32 mode is expected for DTrace traps. PR: https://bugs.freebsd.org/bugzilla/show_bug.cgi?id=298064 MFC after: 1 month Reviewed by: markj Differential Revision: https://reviews.freebsd.org/D59279
bus_dmamap_load_mem() reports most mapping failures, including EFBIG, only through its callback and then returns zero. nvme_payload_map() logged the error without telling the submission path, so the tracker stayed on the outstanding list with no command submitted and no timeout armed, stalling all later I/O on the queue behind it. Approved by: ngie (co-mentor) MFC after: 1 week Reviewed by: ngie, imp Differential Revision: https://reviews.freebsd.org/D59151
The namespace character device does not initialize si_iosize_max, so physio falls back to DFLTPHYS and can produce a bio larger than the qpair payload DMA tag on a controller whose maximum transfer size is below 64KB. Such a bio fails DMA mapping and is never submitted. Approved by: ngie (co-mentor) MFC after: 1 week Reviewed by: ngie, imp Differential Revision: https://reviews.freebsd.org/D59152
This change adds support for AMD's UMC performance counters. It is a bit more complicated than existing counters because the enable bit has moved. This supports Zen 4 through most Zen 6 chips as UMC counters are per-node, where a node does not necessarily translate to a NUMA domain. A few follow up changes to PMC will address this limitation. Reviewed by: mhorne Sponsored by: Netflix Pull Request: https://github.com/freebsd/freebsd-src/pull/2368
The initial statistics update runs before the PF VSI has obtained its firmware-assigned statistics counter index. Discard that provisional VSI baseline so the first update after initialization records the correct hardware counter. Without this reset, subtracting a larger provisional value from a newly selected counter can be mistaken for a 32-bit wrap and report nearly UINT32_MAX receive drops immediately after boot. Reported by: Daniel Braniss <danny@cs.huji.ac.il> Tested by: Daniel Braniss <danny@cs.huji.ac.il> Obtained from: Intel ixl 1.14.2 MFC after: 2 weeks Sponsored by: BBOX.io Differential Revision: https://reviews.freebsd.org/D59336
Whenever we create a user-space mapping, we always set ATTR_S1_PXN in the PTE, which blocks execution of user-space code while running in kernel mode. However, when seeking to determine whether we need to perform an icache flush before installing the new PTE, we test whether sometimes the old PTE or other times the new PTE has ATTR_S1_XN set. The trouble is that ATTR_S1_XN is defined as the bitwise OR of ATTR_S1_PXN and ATTR_S1_UXN, and so the test for whether ATTR_S1_XN is set is satisfied if either of its constituent bits is set, i.e., we write (l3e & ATTR_S1_XN) != 0. Consequently, the test is always true. In practice, I believe that the ill effects of this bug are limited: In pmap_enter(), in rare circumstances, e.g., wiring a code page, an unnecessary icache flush will be performed. In pmap_enter_l2() and pmap_enter_l3c(), no icache flush will be performed. However, typically an icache flush would have already been performed on each of the constituent base pages. Reviewed by: kib, markj MFC after: 3 weeks Differential Revision: https://reviews.freebsd.org/D59265
Wake-on-LAN capability was inferred from NVM bits on every MAC even though 82599 support is board and sometimes port specific. Private sysctls formed a second policy interface, and the driver neither coordinated the controller wake source with PCI PME nor reliably rebuilt address filters erased by the stop-time reset. Use the standard ifconfig wake capabilities. Derive support from the 82599 board and port matrix or the X540-and-newer NVM capability. Require D3hot PME support, and use the NVM APME bit only to select the initial magic-packet policy after initializing the LAN function number. Snapshot requested filters before the terminal stop so shared reset and PHY code sees the active wake policy. After reset, restore RAR0, the multicast table, receive filtering, and the optical laser before arming WUFC, WUC, and PCI PME. Remove device wake sources before clearing PCI PME on detach, resume, and when wake is disabled. Clear autonomous APM so ifconfig remains authoritative. Treat X550EM low-power-link-up failure as best effort and allow shutdown to continue after a wake-programming error. The 82599, X540, X550, and E610 datasheets document the standard ACPI wake filters used here; the E610 ACPI path includes magic-packet wake. Validated on a dual-port E610. Both ports completed three direct-D3 cycles covering wake disabled and magic-packet wake armed. A system S3 cycle woke through ix0 with WUS 0x00000002 (magic packet). Link and traffic recovered after each transition. Note that many add-in cards in this family do not support WoL; LOM and OCP cards are more likely. The E610 as tested does. MFC after: 2 weeks Sponsored by: Dirk-Willem van Gulik from Web Weaving (E610 hardware) Sponsored by: BBOX.io
Reported by: Andrew Griffiths <andrew@calif.io> Reported by: Chris Jarrett-Davies <chrisjd@openai.com> Sponsored by: The FreeBSD Foundation MFC after: 1 week
There are 4 options for e500 watchdog timeout, which may be core- or even SoC- specific. Add support to tune the behavior via a tunable (machdep.watchdog_mode). The tunable value is an integer 0-3.
A critical exception, such as a watchdog, can trigger at any time, including the middle of a standard exception prologue or epilogue, so GPRs, including %r1 (the stack pointer) cannot be trusted at all. Instead, use a private stack pointer for critical interrupts. Each CPU now has its own critical exception stack, with the boot stack in the bss.
There's a small window between when the SRR* registers are restore and the exception returns, in which a TLB miss exception may be triggered. Since there are not special SRR* registers for TLB miss exceptions, the registers from the frame will be ovwritten, and the FRAME_LEAVE block will effectively be re-entered on exit, leading to a very hard to diagnose panic or wedge. Minimize this chance by pushing the SRR* restore to the last possible moments, caching them in a PCPU save area instead until the end. This matches what the AIM side already does.
If a pmap is freed and its memory is reused before its TID reference is taken, then arbitrary memory will be clobbered. Avoid this by never dereferencing the pmap pointer in the tidbusy array, and instead using it as a compare sentinel.
When a TLB miss exception occurs the exception handler must walk the page table from the root. When the root is allocated from KVA the TLB miss exception may take another exception if the root page(s) aren't in the TLB. The 64-bit page table is modeled after the AIM radix page table, with a 64kB root "page", so 16 pages. This makes regular use of UMA allocations unable to refer back to the DMAP, which itself is mapped in TLB1. Now we take another page from the radix pmap driver and grab contiguous pages from the VM system, so that we can simply refer directly to DMAP and avoid more nested TLB misses. We can still take a nested miss, though, because the pmap itself may be in KVA, but this reduces the nesting.
Add a new CPU-family `show pcpu` handler, cpu_db_show_mdpcpu() to dump CPU-specific PCPU data. Only Book-E is populated for now, but AIM may be populated later. These new field prints: save areas, TLB miss nesting, the new critical stack pointer. All of them have been very useful for debugging very esoteric bugs, so make them easier to see from DDB, instead of having to rummage through hex dumps.
The initializer just needs working malloc(9) and two constants that are set at hammer_time(). Fixes: https://cgit.freebsd.org/src/commit/?id=648fa3558c161a1d8564626d21047710c3fbfdf6 Reviewed by: avg, markj Differential Revision: https://reviews.freebsd.org/D58714
When debugging the D59463 review, it is very handy to be able to change the software portal holdoff time without recompiling the kernel. This commit makes the holdoff time a sysctl tunable, so it can be changed at runtime. Tested by: dsl Obtained from: flo_purplekraken.com MFC after: 3 weeks Differential Revision: https://reviews.freebsd.org/D59461 Event: Berlin Hackathon 202609
Expose cached per-VF configuration through the iflib VF status method. Report mailbox handshake state, MAC address, access or trunk VLAN mode, hardware transmit and receive queue counts, administrator policy, and fault-containment state. The query does not issue mailbox requests or read hardware registers. Sponsored by: BBOX.io
Expose cached VF configuration, policy, and runtime state through the iflib VF status method. Include access or trunk VLAN mode, transmit and receive queue counts selected by the current virtualization mode, negotiated mailbox API, PF traffic permission, fault containment, and quarantine state. Initialize every cached API version before VF enumeration so an unconfigured slot cannot be mistaken for API 1.0. The query does not issue mailbox requests or read hardware registers. Sponsored by: BBOX.io
Expose cached per-VF configuration through the iflib VF status method. Report mailbox initialization and the negotiated virtual-channel API, MAC address, access or trunk VLAN mode, queue resources, administrator policy, PF traffic permission, and fault containment. Expose per-VF malicious-driver isolation and cumulative transmit and receive event counts through a versioned driver.ixl extension. Track successful PCI IOV attachment separately from hardware capability. This lets a successfully attached but unconfigured PF return an empty snapshot without claiming support when PCI IOV registration was unavailable. The query uses driver-cached state and does not issue AdminQ requests or read device registers. Sponsored by: BBOX.io
Add dev.asmc.0.sil sysctl to control the SIL LED via SMC keys MSLD (duty/brightness) and MSLS (state latch, must be set before re-enabling). MFC After: 1 week Reviewed by: adrian Differential Revision: https://reviews.freebsd.org/D58865
apple_bce: kick USB explore thread after VHCI attach Call usb_needs_explore() after attach so the hub explore thread enumerates all initially connected ports instead of only the first. Reviewed by: adrian Differential Revision: https://reviews.freebsd.org/D58870
apple_bce: fix cold boot panic in mailbox send Poll mailbox reply registers with DELAY() when the system is still cold, falling back to the interrupt-driven sema_timedwait path once timers are available. Reviewed by: adrian Differential Revision: https://reviews.freebsd.org/D58871
apple_bce: ignore duplicate TRANSFER_REQUEST in STATUS state Since the host initiates IN data transfers, the firmware's own TRANSFER_REQUEST for the same phase arrives after we've already moved to STATUS state. Ignoring instead of failing the transfer with USB_ERR_IOERROR. Reviewed by: adrian Differential Revision: https://reviews.freebsd.org/D58873
apple_bce: fix C_CONNECT_STATUS during port reset Only set the port change bit when the corresponding status bit actually transitioned, instead of unconditionally flagging C_CONNECT_STATUS on every port change event. Also clear any flaky C_CONNECT_STATUS that the port change taskqueue may have set while the bus lock was dropped during a successful port reset. Reviewed by: adrian Differential Revision: https://reviews.freebsd.org/D58875
apple_bce: initiate IN data phase for control transfers For IN control transfers, the firmware does not send TRANSFER_REQUEST for the data phase the host must send BCE_VHCI_CMD_TRANSFER_REQUEST with an IN DMA buffer once the setup phase completes. Reviewed by: adrian Differential Revision: https://reviews.freebsd.org/D58872
The shared code already recognizes IAVF_DEV_ID_VF_HV and handles it through the regular iavf register and virtchnl paths, but the PCI probe table omits it. Add the missing entry so the driver attaches. PR: https://bugs.freebsd.org/bugzilla/show_bug.cgi?id=239849 MFC after: 1 week
iflib releases queue mappings after IFDI_STOP() returns. If a Physical Function reset fails, returning with PCI bus mastering enabled can therefore leave the device able to access mappings which the framework is about to recycle. Disable bus mastering and drain pending PCIe transactions when reset fails during initialization or stop. Re-enable and verify bus mastering only after a later reset succeeds and before queue programming begins. Treat inability to establish the DMA fence as a fail-stop invariant violation. MFC after: 2 weeks Sponsored by: BBOX.io
dsp_chn_alloc() stopped at the first primary channel that was either idle or already had vchans. Since the list is walked in order, the first channel matched both conditions once it had been used, so every client after the first was stacked onto it as a vchan and the remaining primary channels were never allocated at all. This is invisible on devices with a single primary channel, but not on those which provide several. snd_emu10kx(4), for instance, registers four primary channels for its front device, each able to run with its own rate. Look for an idle primary channel first, and only fall back to sharing one that already has vchans when there is none left. Sponsored by: The FreeBSD Foundation MFC after: 2 weeks Reviewed by: markj Differential Revision: https://reviews.freebsd.org/D59084
Makes it easier to test scenarios involving more than 1 primary channels per direction. Sponsored by: The FreeBSD Foundation MFC after: 2 weeks Reviewed by: markj Differential Revision: https://reviews.freebsd.org/D59085
The playback voices always loop over the whole EMU_PLAY_BUFSZ buffer, but emupchan_setblocksize() only recorded the new block size and left the block count as it was set up by emu_vinit(). The blocks then no longer covered the whole buffer, and the part they left out was played without ever being written to, which became audible as distortion once playback started going through a virtual channel. Resize the buffer, so that the block count and size always cover it. Fixes: https://cgit.freebsd.org/src/commit/?id=02d4eeabfd73 ("sound: Allocate vchans on-demand") PR: https://bugs.freebsd.org/bugzilla/show_bug.cgi?id=287687 MFC after: 1 week Sponsored by: The FreeBSD Foundation Reviewed by: markj Differential Revision: https://reviews.freebsd.org/D59444
We don't consistently clear PGA_WRITEABLE on fictitious, managed pages. Some functions do, e.g., pmap_remove_all(), but several do not. At worst, this is just a pessimization, but there is no good reason to be inconsistent. Clear PGA_WRITEABLE in those that previously did not by introducing (and using) the helper function pmap_page_is_mapped_locked() that implements the correct test. Reviewed by: kib, markj MFC after: 3 weeks Differential Revision: https://reviews.freebsd.org/D59466
create_rss_raw_qp_tir() wrote MLX5_RX_HASH_FUNC_TOEPLITZ (the userspace ABI flag, value 1) into tirc.rx_hash_fn. That field takes the hardware encoding from mlx5_ifc.h, where 1 is INVERTED_XOR8 and Toeplitz is MLX5_TIRC_RX_HASH_FN_HASH_TOEPLITZ (2). Reviewed by: kib, slavash Sponsored by: Nvidia Networking MFC after: 1 week
USB Vendor:Product 0x2357:0x0115 The data was provided by Pavel Timofeev (timp87 gmail com) and the device will be supported by rtw88 USB (rtw8822bu.c) in the future. MFC after: 3 days
bce_vhci_endpoint_create() was registering both the IN and OUT submission queues on every call, regardless of which direction the endpoint actually needed. The per-stage firmware registration failures weren't rate-limited, so a stuck registration will spam the console. Tested on: MacBookPro15,2 Macmini8,1 MacBookPro15,3 MacBookPro16,2 Reviewed by: adrian Differential Revision: https://reviews.freebsd.org/D58831
The Create I/O queue commands carried magic 0x1/0x3 cdw11 values, and the implicit QPRIO of zero classifies every I/O queue Urgent the moment weighted round robin arbitration is enabled. No functional change with the round robin default. Approved by: ngie (co-mentor) Reviewed by: ngie, imp, adrian Differential Revision: https://reviews.freebsd.org/D59153
MAXCMD is the maximum number of commands the controller processes at one time for a particular I/O queue. Allocating more trackers than that only queues commands inside the controller. Controllers reporting 0 are unaffected. Approved by: ngie (co-mentor) Reviewed by: ngie, imp Differential Revision: https://reviews.freebsd.org/D59154
The original implementation assumed that the start/reply doorbells lived within the device _CRS space, but that isn't always the case. On my AMD Ryzen 7640U-based frame.work laptop, device memory runs from 0xc0500000-0xc0500fff while the doorbells are up around 0xc0508000. Stop sanity checking the addresses and just map them in to work reliably whether they're within the device range or not. pluton_wait_reply is cribbed from tpm_wait_for_u32, but rewritten slightly to read in just one place and to read one last time before giving up at the end of the timeout, just in case. Reviewed by: kbowling Differential Revision: https://reviews.freebsd.org/D59327
arm64: Fix the CTR_EL0 field names Use the field names from the Arm Documentation for the CTR_EL0 register. These will later be generated from a BSD Licensed JSON file so to reduce the diff for that rename now. Reviewed by: Kajetan Puchalski <kajetan.puchalski@arm.com> Sponsored by: Arm Ltd Differential Revision: https://reviews.freebsd.org/D59173
arm64: Fix the ICC_SGI1R_EL1 field names Use the field names from the Arm Documentation for the ICC_SGI1R_EL1 register. These will later be generated from a BSD Licensed JSON file so to reduce the diff for that rename now. Reviewed by: Kajetan Puchalski <kajetan.puchalski@arm.com> Sponsored by: Arm Ltd Differential Revision: https://reviews.freebsd.org/D59174
arm64: Fix the ID_AA64MMFR1_EL1.CMOW name This was incorrectly names CMOVW. Remove the unneeded 'V'. Reviewed by: Kajetan Puchalski <kajetan.puchalski@arm.com> Sponsored by: Arm Ltd Differential Revision: https://reviews.freebsd.org/D59175
arm64: Remove ID_AA64PFR0_GIC_BITS
It's unused since 05f6f65c3bda ("arm64: add CHECK_CPU_FEAT() for
checking feature support in assembly"). It's been replaces with a
commonly named ID_AA64PFR0_GIC_WIDTH macro.
Reviewed by: Kajetan Puchalski <kajetan.puchalski@arm.com>
Sponsored by: Arm Ltd
Differential Revision: https://reviews.freebsd.org/D59177
arm64: Rename some ID_AA64DFR0_EL1 macros These macros are used to extract the number of breakpoints and watchpoints from the register. Rename them to not conflict with the common meaning of an _VAL macro in this file. Reviewed by: Kajetan Puchalski <kajetan.puchalski@arm.com> Sponsored by: Arm Ltd Differential Revision: https://reviews.freebsd.org/D59178
arm64: Add ICC_SGI1R_IRM_VAL and use it This will be the name of the generated macro. As we only need to check if this field is set or not use the common ICC_SGI1R_IRM_VAL name. Reviewed by: Kajetan Puchalski <kajetan.puchalski@arm.com> Sponsored by: Arm Ltd Differential Revision: https://reviews.freebsd.org/D59179
arm64: Fix the capitalization of CurrentEL Use the same capitalization as the Arm documentation. Reviewed by: Kajetan Puchalski <kajetan.puchalski@arm.com> Sponsored by: Arm Ltd Differential Revision: https://reviews.freebsd.org/D59180
arm64: Move the CNTV_CTL_EL0 fields to armreg.h They are architecturally defined, so should be in the common location. Reviewed by: emaste, Kajetan Puchalski <kajetan.puchalski@arm.com> Sponsored by: Arm Ltd Differential Revision: https://reviews.freebsd.org/D59181
ICH_VTR_EL2.PREbits contains the number of virtual preemption bits implemented minus one, analogous to ICH_VTR_EL2.PRIbits. Unlike PRIbits, the macro for PREbits does not account for "minus one". Make the two match. Signed-off-by: Kajetan Puchalski <kajetan.puchalski@arm.com> Reviewed by: andrew Sponsored by: Arm Ltd Pull Request: https://github.com/freebsd/freebsd-src/pull/2391
Fix incorrect opcode on CLIDR_EL1. Signed-off-by: Kajetan Puchalski <kajetan.puchalski@arm.com> Reviewed by: andrew Sponsored by: Arm Ltd Pull Request: https://github.com/freebsd/freebsd-src/pull/2400
pmap_stage2_fault is supposed to invalidate the icache if it is accessing an executable page. Instead, the condition currently checks whether execution is disallowed at any EL. Make it check whether execution is allowed at any EL to match the intended behaviour. Signed-off-by: Kajetan Puchalski <kajetan.puchalski@arm.com> Reviewed by: andrew Sponsored by: Arm Ltd Pull Request: https://github.com/freebsd/freebsd-src/pull/2416
Reviewed by: andrew Sponsored by: Arm Ltd Differential Revision: https://reviews.freebsd.org/D59170
Reviewed by: andrew Sponsored by: Arm Ltd Differential Revision: https://reviews.freebsd.org/D59168
Reported by: alc Tested by: alc Reviewed by: andrew Sponsored by: Arm Ltd Differential Revision: https://reviews.freebsd.org/D59476
Add support for reading the MPERF (MSR 0xE7) and APERF (MSR 0xE8) model-specific registers on AMD/Intel CPUs through hwpmc(4). These counters track maximum and actual performance frequency respectively, and are used to compute effective CPU frequency scaling independent of the nominal TSC rate. The name of the class was chosen as PERF because later support for other PERF MSRs can be added to the same class. Extend libpmc(3) to expose the AMD/Intel MPERF/APERF counters added to hwpmc(4) in the companion kernel change, so userland consumers (pmcstat(8), etc.) can allocate and read these events by name. Document the new PERF class and its MPERF/APERF counters in a new pmc.perf.3 manual page, describing their semantics and how to read them via pmc(3) and pmcstat(8). Bump PMC_VERSION_MINOR. Signed-off-by: Anderson Nascimento <anascime@amd.com> Reviewed by: Ali Mashtizadeh <ali@mashtizadeh.com>, mhorne Reviewed by: ziaee (manpages) Sponsored by: AMD Differential Revision: https://reviews.freebsd.org/D58647
Previously each channel allocated its own set of buffers using the channel's DMA tag which causes a DMA lock contention under load. Even a single saturated 1 Gbps link caused ~50,000 adaptive mutex spin events (as per lockstat) per second. With the proposed approach there's no more contention on the channel's DMA mutex and the adaptive mutex spin events dropped to ~6,000/s. Stress test where iperf3 pushed as much traffic as possible to the 4 ports revealed that throughput drops from expected 940 Mbps down to 600-800 on each link with the "bounce pages lock" generating ~110,000 adaptive mutex spin events per second, but this is to be addressed later on. Tested by: flo_purplekraken.com, gnikl_justmail.de, bofh@ MFC after: 3 weeks Differential Revision: reviews.freebsd.org/D59463 Event: EuroBSDcon Devsummit 2026
Reviewed by: dsl Approved by: dsl Obtained from: flo_purplekraken.com MFC after: 3 weeks Differential Revision: https://reviews.freebsd.org/D59497 Event: EuroBSDcon Devsummit 2026
This fixes the build with gcc 16 after the changes to -Wunused*[0]. [0] https://gcc.gnu.org/gcc-16/porting_to.html#changes-to-wunused Reviewed by: np MFC after: 3 days Sponsored by: The FreeBSD Foundation Differential Revision: https://reviews.freebsd.org/D59545
Commit 851dffef532a added the Lunar Lake-M I2C controllers 0 through 3 (0xa878-0xa87b), which sit on PCI device 0x15. The platform exposes two further controllers at 0xa850 and 0xa851 on PCI device 0x19, reported by Intel as I2C #4 and #5. This mirrors the layout already handled for Arrow Lake-U, where both the 0x777x and 0x775x ranges are listed. On an HP OmniBook X Flip 16-as0xxx (Core Ultra 9 288V) the firmware enables only four of the six controllers, and both HID devices sit on the two that were missing: an ELAN2514 touchscreen on controller 4 and a SYNA3503 touchpad on controller 5. Neither attaches without this change, so the machine has no working pointing device. Like the other four, these use the Tiger Lake revision of the I2C IP; Linux treats 0xa850/0xa851 identically to 0xa878-0xa87b in intel-lpss-pci.c. Tested on: HP OmniBook X Flip 16-as0xxx (Intel Core Ultra 9 288V) Reviewed by: vexeduxr MFC after: 2 weeks Signed-off-by: Kang Kang <kk1987@gmail.com> Pull Request: https://github.com/freebsd/freebsd-src/pull/2406
If we skip one of the constraint package elements and decrement sc->constraint_count, sc->constraints would end up sparse and some of our constraints would be put after sc->constraint_count. Reviewed by: olce Sponsored by: The FreeBSD Foundation Event: EuroBSDCon Devsummit 2026 Differential Revision: https://reviews.freebsd.org/D59562
AMD doesn't encode page depth in eptp. As a result, the page level is decided by the host la57 value. Without this, it uses 4 level page and therefore cause machine enable la57 have garbage page translation. PR: https://bugs.freebsd.org/bugzilla/show_bug.cgi?id=291768 Reviewed by: kib Tested by: Antranig Vartanian <antranigv@freebsd.am> MFC after: 1 week Differential Revision: https://reviews.freebsd.org/D57978
Add a driver for the Sony PS5 DualSense controller (054c:0ce6) providing gamepad input via evdev, lightbar RGB LED control, and player indicator LEDs through sysctl. Signed-off-by: Christos Longros <chris.longros@gmail.com> Reviewed by: ziaee, wulf MFC after: 1 month Differential Revision: https://reviews.freebsd.org/D56345
The firmware recovery callout continues polling controller registers after a power transition. An inaccessible E610 GL_MNG_FWSM register reads as all ones in D3, which looks like firmware recovery mode and queues an iflib reset while the device is suspended. A later D0 poll then reports recovery complete and queues another reset. Pause and drain the callout before terminal stop policy is applied, prevent an in-flight callback from rearming it, and restart polling only after resume has cleared the wake state. Track callout initialization so partial attach cleanup does not drain an uninitialized callout. The false transition was reproduced on a dual-port E610 with direct D3 and system S3. Validate the guard with wake-disabled and wake-armed D3, three repeated D3 cycles per port, and an S3 magic-packet wake. Both ports returned to D0 without a false recovery transition. MFC after: 2 weeks Sponsored by: BBOX.io
Do not let iflib publish a running interface when the virtual device rejected its enable command. Mark initialization failed and leave the interface stopped. MFC after: 2 weeks Sponsored by: BBOX.io
ixl_set_state() takes a bit index, not a bit mask. ORing the reset request and critical error indices produced the global reset index, so a critical interrupt could mask its cause without scheduling the intended PF reset. Set both state bits explicitly. MFC after: 2 weeks Sponsored by: BBOX.io
PF iflib callbacks receive struct ice_softc, not struct ice_mirr_if. Resolve the mirror interface through sc->mirr_if before checking or resetting subinterface state. Use the same mirror softc when rebuilding its VSI. This records the required subinterface reset in the state consumed by the PF callback, rather than overlaying the PF softc and leaving rebuilt queues stopped. MFC after: 2 weeks Sponsored by: BBOX.io
Rebuilding the Flow Director tables does not initialize the interface or restore its queues. Do not set IFF_DRV_RUNNING from that operation: iflib owns the flag, and may have cleared it while a watchdog reset is pending. Restoring it here could admit traffic before the deferred stop and initialization have run. Keep the table rebuild and Flow Director interrupt re-enable unchanged. This path is conditional on IXGBE_FDIR. MFC after: 2 weeks Sponsored by: BBOX.io
Add a new "high temp" threshold below the max temp, in order to ramp the fans (or pump for liquid cooled quads) sooner, and avoid the maximum temperature. From the PR, this makes older G5 quads louder, but usable, instead of hitting max temperature and forcing a thermal shutdown. PR: https://bugs.freebsd.org/bugzilla/show_bug.cgi?id=209202 Submitted by: gmbroome (PR)
The dongle uses the rtl8723bu chip. Also add an entry for it in usbdevs.
No functional change, just simplify the code flow a little.
Machine check and critical exceptions are independent of each other, and can interrupt each other. Since they're asynchronous we cannot trust that the existing stack pointer (%r1) is correct at time of entry, so add a private machine check stack separate from the critical exception stack. As part of this, switch to the STANDARD_CRIT_PROLOG() for machine check exceptions, and overload the macro to specify the stack to switch to. Also add a savearea argument to CRIT_SRR_RESTORE() so that we can restore machine check exception state from the right location.
Some firmware inexplicably decides to do non-standard and annoying stuff
here, e.g. the Fujitsu Futro S940 with an Intel Pentium J5005 sometimes
returns the following when calling the DEVICE_CONSTRAINTS function on
the Intel DSM:
Return (Package (0x01)
{
Zero
})
(Package elements here are supposed to be constraint packages, not just a
single value.)
First reported in the following forum post:
https://forum.netgate.com/topic/201090/2.9.0-beta-leads-to-kernel-panic-on-boot
Reported by: TampertK on forum.netgate.com
Reviewed by: olce
Sponsored by: The FreeBSD Foundation
Event: EuroBSDCon Devsummit 2026
Differential Revision: https://reviews.freebsd.org/D59566
Also match behaviour with Intel constraint parsing by skipping malformed constraints instead of failing hard. Reviewed by: olce Sponsored by: The FreeBSD Foundation Event: EuroBSDCon Devsummit 2026 Differential Revision: https://reviews.freebsd.org/D59568
MMUCFG::PIDSIZE is the TID field size less 1, so adjust to get the full width.
QEMU doesn't appear to emulate the MMUCFG register for Book-E CPUs, so give a sane small default of 7, for an 8 bit PID register.
The ASUS PRIME X570-P carries a Nuvoton NCT6798D, Super I/O device ID 0xd42b. Add an exact-match entry to both. Exact rather than masked: the neighboring 0xd42a entries are deliberately exact with an extid because that ID is claimed by both NCT6796D-E and NCT5585D, and widening the family would make them collide. The NCT6798D has seven tachometers, so raise NCTHWM_FAN_MAX to seven and describe the two extra ones; existing entries keep fan_count = 5 and are unaffected. Fan names follow the NCT6779 convention and do not map to any board's physical headers. Tested on: ASUS PRIME X570-P, Ryzen 9 5950X, FreeBSD 16.0-CURRENT. Approved by: adrian Reviewed by: stephane.rochoy_stormshield.eu, adrian Differential Revision: https://reviews.freebsd.org/D58291 Signed-off-by: Nick Price <nprice@FreeBSD.org>
puc has preferred MSI for every card since MSI support was added, with only a global tunable to opt out. uart(4) makes the same decision for the serial devices it attaches directly, and has since grown two defences: it skips MSI unless the device advertises exactly one vector, because attaching a single instance to a device offering many has caused problems (PR 235016), and it lets individual devices be flagged when they claim MSI support that does not work. Adopt both. Approved by: adrian (mentor) Reviewed by: adrian Differential Revision: https://reviews.freebsd.org/D59623 Signed-off-by: Nick Price <nprice@FreeBSD.org>
Two 16850 UARTs in the first I/O BAR at offset 0xc0, 8 bytes apart. Approved by: adrian (mentor) Reviewed by: imp, adrian Differential Revision: https://reviews.freebsd.org/D59512 Signed-off-by: Nick Price <nprice@FreeBSD.org>
tpm(4) was removed from amd64 GENERIC because it broke suspend and resume. The preceding lifecycle, state-save, interrupt, locality, and teardown fixes address those failures for both TPM 1.2 and TPM 2.0. Restore the driver to amd64 GENERIC and MINIMAL, where TPM entropy harvesting remained enabled. Enable the driver and entropy harvesting in the MPC85XX and QORIQ64 configurations, which already provide FDT, spibus, and the platform SPI controller required by FDT-attached TPMs. Leave the generic AIM and POWER configurations unchanged because they have no TPM attachment bus. The TPM 1.2 path completed repeated S3 cycles and command tests on ThinkPad T430 and T440p systems. The TPM 2.0 path completed repeated device and full-system suspend/resume cycles on a ThinkPad P51. The PowerPC configuration matrix was checked to retain tpm(4) only where its FDT SPI attachment path is present. PR: https://bugs.freebsd.org/bugzilla/show_bug.cgi?id=291067 Reviewed by: kevans Tested by: Marek Zarychta <zarychtam_plan-b.pwste.edu.pl> (tpm1.2) Fixes: https://cgit.freebsd.org/src/commit/?id=16f8ea6a81b5 ("amd64: Remove tpm(4) from GENERIC for now") MFC after: 1 month Relnotes: yes Sponsored by: BBOX.io Differential Revision: https://reviews.freebsd.org/D59247
move thermal shutdown arming to the admin poll this gives a more reasonable delay prior to the first attempt, and also allows us to retry and make the option runtime-tuneable via a new disable_thermal_arm sysctl Approved by: adrian (mentor) Reviewed by: adrian Differential Revision: https://reviews.freebsd.org/D59593 Signed-off-by: Nick Price <nprice@FreeBSD.org>
When interface is up and running 'kldunload if_iwx' stops the device and executes RUN -> INIT state transition. Since the device is already stopped iwx_run_stop fails to stop the device again and returns non-zero exit code from iv_newstate callback which triggers 'INIT state change failed' assertion. I reused IWX_FLAG_SHUTDOWN flag to: a) set it in iwx_detach b) check it in iwx_newstate_sub - when it is set all custom state transition logic is skipped Accidentally found while experimenting with iwlwifi / iwx drivers Reviewed by: adrian Differential Revision: https://reviews.freebsd.org/D59624
Chromebook keyboards driven by the ChromeOS EC emit the top-row action keys as scancode set 1 codes 0xE0 0x11 - 0xE0 0x1E (see the codeset 1 comments on the SCANCODE_* enum in chrome-ec include/keyboard_8042_sharedlib.h). The e0 half of the evdev_scancode2key() translation table leaves eleven of those at NONE, so on FreeBSD those keys emit no evdev event at all and userspace remappers such as keyd have nothing to bind to: e0 11 fullscreen KEY_ZOOM e0 12 overview KEY_SCALE e0 13 screenshot KEY_SYSRQ e0 14 brightness down KEY_BRIGHTNESSDOWN e0 15 brightness up KEY_BRIGHTNESSUP e0 16 privacy screen toggle KEY_PRIVACY_SCREEN_TOGGLE e0 17 kbd backlight down KEY_KBDILLUMDOWN e0 18 kbd backlight up KEY_KBDILLUMUP e0 1a play/pause KEY_PLAYPAUSE e0 1b mic mute KEY_MICMUTE e0 1e kbd backlight toggle KEY_KBDILLUMTOGGLE Fill those in. Only entries that were previously NONE are touched: the EC's remaining top-row codes already have mappings here (e0 10 prev track, e0 19 next track, e0 20 mute, e0 2e / e0 30 volume, e0 67 refresh, e0 69 forward, e0 6a back), as do e0 1c keypad enter and e0 1d right control, and none of them change. Signed-off-by: Kang Kang <kk1987@gmail.com> Tested on: Lenovo ThinkPad C14 Chromebook (Google primus, ChromeOS EC) Reviewed by: wulf MFC after: 2 weeks
The current interrupt path uses the IRQ resource value directly as the LPC SIRQ selector in TPM_INT_VECTOR and already restricts it to 1 through 15. This is a driver limitation: a parent interrupt number need not equal an LPC SIRQ channel, and SPI TPMs can use a separate parallel interrupt. On the reported system with ACPI IRQ 45, the existing range check runs after handler registration and returns before disabling firmware interrupt delivery. This can leave a polling device with a handler on an asserted source. Disable and verify interrupt delivery before registering a handler or starting common TPM services. Preserve the existing range policy, using polling without registering a handler for routes rejected by that check, and release their IRQ resources. Keep a failed setup's potentially stale output cookie out of the device state; the interrupt framework may already have removed that handler. Cancel timed-out locality requests in the shared request helper, covering both interrupt programming and commands. Permit attach after a locality timeout if the enable register proves delivery is already off, allowing later command recovery. If delivery remains enabled and cannot be quiesced, fail attach instead of exposing a device node for possible later recovery. Report this as a quiesce failure since locality acquisition as well as register programming can fail. On detach, publish dying and serialize with commands before quiescing TPM interrupt delivery. Do this for polling devices too, since firmware may have re-enabled delivery and resume-time quiescing may have failed. Quiesce before common release destroys the command lock and before removing any handler. Report hardware quiesce failure while completing software cleanup. Reported by: adrian Reviewed by: adrian, imp MFC after: 2 weeks Sponsored by: BBOX.io Differential Revision: https://reviews.freebsd.org/D59645
Add driver for Intel Power Management Controller (PMC) found on Sunrise Point PCH chipsets. This device exposes S0ix sleep state residency counters and power management status. Sysctls provided: dev.intelpmc.0.slp_s0_residency_us - Time in deepest sleep (us) dev.intelpmc.0.ltr_ignore - LTR ignore mask dev.intelpmc.0.pm_cfg - PM configuration register dev.intelpmc.0.pm_sts - PM status register dev.intelpmc.0.access_denied - Firmware lock status Supported devices for now: - Sunrise Point-LP (0x9D21) - Sunrise Point-H (0xA121) Note: Later PCH generations (Cannon Lake, Tiger Lake, ect.) have different PMC register layouts according to the datasheet and would need per-generation tables. I avoided adding untested hardware in case they have a quirk. Reviewed by: olce, adrian Differential Revision: https://reviews.freebsd.org/D54881
Firmware expose a DSDT sized for the largest SKU of the platform, so a verbose boot prints an "ignored" line for every vacant processor slot. A vacant slot has no enabled MADT entry; a CPU that failed to come online does. Reviewed by: olce, adrian Differential Revision: https://reviews.freebsd.org/D59551
this fixes, 2 failed v4l2-compliances tests
selwakeup() only wakes select/poll waiters; it does not notify the kqueue knote registered on the device, so EVFILT_READ never fired when a frame became available.
mmu_radix_extract() walks the page tables without holding the pmap lock, unlike its hash MMU counterpart moea64_extract(). A concurrent unmap can free and recycle the page table page being walked, so the read returns whatever now occupies that memory and the caller gets a physical address that never existed. That is how mmu_radix_sync_icache() came to hand a bogus address to __syncicache() and panic the machine. Commit 1574ca1955f5 worked around it by taking the pmap lock in mmu_radix_sync_icache(), but the machine independent callers of pmap_extract() - vm_sync_icache(), proc_rwmem() and the vslock() paths - remain exposed to the same failure. Rename the existing body to mmu_radix_extract_locked(), which asserts the lock, and make mmu_radix_extract() a thin wrapper that acquires it. mmu_radix_sync_icache() already holds the pmap lock, so it calls the locked variant directly and neither recurses nor reacquires the lock once per page. Suggested by: alc MFC after: 1 week Differential Revision: https://reviews.freebsd.org/D59320 Reviewed by: markj, jhibbits
tpmcrb_transmit() read CRB_CTRL_STS before requesting locality 0. An AMD Pluton fTPM (like FrameWork Desktop) using the plain CRB start method reads the control area as all-ones until locality is assigned, so bit 0 looks like a stuck tpmSts and every command failed with EIO. With RANDOM_ENABLE_TPM the harvester retries every 10 seconds, so this printed "Device has Error bit set" forever. Reviewed by: kbowling Approved by: kbowling MFC after: 2 weeks Sponsored by: Netflix Differential Revision: https://reviews.freebsd.org/D59661
Always offer the VIRTIO_NET_F_GUEST_CSUM feature to the host, and not only if RXCSUM is enabled on vtnet. Instead of using RXCSUM to control whether this feature will be negotiated with the host, just use it to control whether the VIRTIO_NET_HDR_F_DATA_VALID flag on an incoming packet is processed (i.e., translated to the corresponding mbuf flag only if RXCSUM is enabled on the vtnet interface). This has two benefits: 1. Enabling/disabling RXCSUM on vtnet does not require feature renegotiation. 2. The host is always allowed to send locally generated TCP or UDP packets to the guest without computing a full checksum (by setting the VIRTIO_NET_HDR_F_NEEDS_CSUM flag) and not only if the guest has enabled RXCSUM on vtnet. Since locally generated packets do not require a checksum, this saves otherwise unnecessarily wasted computing power. If a user of a FreeBSD guest really does not want to negotiate the VIRTIO_NET_F_GUEST_CSUM feature with the host, it still can disable the loader tunable hw.vtnet.csum_disable. Reviewed by: kfv, tuexen MFC after: 1 week MFC to: stable/15 Differential Revision: https://reviews.freebsd.org/D59052
VF readiness can precede its first initialization. In particular, iflib populates if_hwassist during init, so copying it while the VF is still down can leave hn advertising checksum and TSO capabilities with no hardware-assist flags. The stack then uses software segmentation until a capability ioctl happens to refresh the copied state. Synchronize enabled capabilities and hardware-assist flags after the VF up ioctl succeeds, before switching the datapath and enabling transparent VF transmit. Use the existing helper in the common initialization path, covering both normal bring-up and delayed VF initialization. MFC after: 2 weeks Sponsored by: BBOX.io
The PCI_BUS_RELATIONS2 case lacks a break after queuing the reported devices. It consequently processes the same packet as PCI_EJECT, interpreting device_count as the slot number. If that value matches an existing child, the callback schedules an unintended device eject. End the relations case after processing the device list, matching the original PCI_BUS_RELATIONS handling. Fixes: https://cgit.freebsd.org/src/commit/?id=ea11861e434a ("arm64: Hyper-V: vPCI: Enabling v-PCI in FreeBSD in ARM64 Hyper-V") MFC after: 2 weeks Sponsored by: BBOX.io
Handle the host's VF association notifications instead of ignoring them. Require an allocated association before switching to the MAC-matched VF, and track notification generations so a withdrawal during initialization cannot enable an obsolete handoff. Defer association and address-event work to the VF taskqueue; the receive channel must remain available to deliver switch completions. Request and wait for the empty VMBus completion for SET_DATAPATH, checking submission and channel-revocation failures. Enable transparent VF transmit and select its link status only after the switch completes. Use the existing transaction lifetime and revocation handling, without a timeout that could leave a late completion referencing a freed request. Restore synthetic capabilities, TSO limits and hardware-assist flags on fallback, targeting hn rather than the departing VF. Block transparent transmit during handoff and after association withdrawal. Separate the attach delay from saved-setting readiness, permit delayed association to trigger initialization, and require a fresh handoff after NVS reattach. Stop and detach still close local VF access if returning to the synthetic path fails; such failures are logged, not reported as a successful switch. MFC after: 2 weeks Sponsored by: BBOX.io
Hold hn_vf_lock across the VF identity/state check and link publication. Otherwise an event can observe an enabled VF, pause during fallback, then overwrite the freshly reported synthetic carrier with stale VF state. Ignore events while switching or after the association generation changes. The link notification only queues further work, so the callback does not need the sleepable hn lock. Document that association serial numbers are currently diagnostic: hn_ismyvf() matches the VF by MAC, while notifications gate availability. MFC after: 2 weeks Sponsored by: BBOX.io
Hyper-V does not expose the normal Intel PF/VF mailbox. Import the Hyper-V operation overrides from DPDK and complete its reset callback with the configuration-space mechanism submitted by Microsoft. Route ixv operations through the operation table, use mailbox API 1.0, and limit this environment to one queue pair. Claim the 82599, X540, X550, X550EM-X, and X550EM-A Hyper-V device IDs listed by DPDK. Select the same path for the E610 Hyper-V subdevice identity defined there. For E610/Linkville, read the emulated VFLINKS-format status from PCI configuration space at offset 0x209. Older families retain the MMIO VFLINKS path. Treat receive-mode changes as host-owned no-ops and do not retry VLAN operations that Hyper-V permanently rejects. Limit the Hyper-V 82599 MTU to 1504 because its VF does not implement the X540 RLPML field. Use mailbox reset indications only to refresh cached Hyper-V link state. Hyper-V does not provide the mailbox handshake required to treat that indication as a request for an iflib reinitialization. Tested on 82599 and E610 (with a followup commit) with Windows Server 2025 Hyper-V. PR: https://bugs.freebsd.org/bugzilla/show_bug.cgi?id=232472, https://bugs.freebsd.org/bugzilla/show_bug.cgi?id=239849 Obtained from: DPDK (shared code) Sponsored by: BBOX.io Co-authored-by: Wei Hu <weh@microsoft.com>
Current Windows PF drivers do not emulate the E610 link-status query at PCI configuration offset 0x209. The installed Intel ixw 1.8.54.0 driver returns zero even with the physical port connected, leaving the VF at no carrier after an otherwise successful hn handoff. When the query returns zero, use VFLINKS for carrier only. Its speed field is not the negotiated E610 line rate: the tested 1 Gb/s port reports the default 10 Gb/s encoding. Publish an unknown speed instead, including in the verbose link-up diagnostic, and document this fallback. Keep nonzero PCI results authoritative and reject all-ones reads. Other Hyper-V families and the native E610 mailbox path are unchanged. A zero result cannot distinguish an unsupported query from a down link; in either case this fallback relies on the legacy carrier indication. The connected and disconnected ports reported VFLINKS.UP set and clear, respectively, while both returned zero from the PCI query. MFC after: 2 weeks Sponsored by: BBOX.io
Discover the Hyper-V queue grant through PCI configuration space for X550, X552, X553, and E610 VFs. The X550 family exposes a byte at 0x207; 0x208 is a separate DCB flag. E610 exposes a little-endian queue count word at those offsets. Accept grants of one, two, or four and use at most two symmetric queue sets, subject to the existing MSI-X resource limit. Keep one queue for invalid grants and for 82599/X540 Hyper-V VFs. The family gates and queue limits are snooped from Windows Intel VF bus traffic. Update the shared code queue max so stop disables every usable queue. Also program the receive length limit on every usable RX queue instead of only queue zero and retain the 82599 exception for that register field. MFC after: 2 weeks Sponsored by: BBOX.io
Some devices only apply a new sample rate once it has been read back, and produce no sound at all otherwise. PR: https://bugs.freebsd.org/bugzilla/show_bug.cgi?id=294803 Reported by: tatsuki_makino@hotmail.com Tested by: tatsuki_makino@hotmail.com MFC after: 2 weeks Sponsored by: The FreeBSD Foundation Reviewed by: emaste Differential Revision: https://reviews.freebsd.org/D59616
parse_mw_buf() copies option names from privileged sysctl input into eight-byte stack buffers. Unbounded %s conversions permit option tokens longer than seven bytes to write past those buffers before the parser validates them. Limit each conversion to seven characters, leaving space for the terminating NUL. Signed-off-by: Yudi Yang <yudi.yang@rice.edu> Fixes: https://cgit.freebsd.org/src/commit/?id=96f556f5044a ("NTB Tool: Test driver for NTB hardware drivers.") Reviewed by: markj MFC after: 1 week
Calling vector_state_store_savectx() overwrites ra with the address of the following ret instruction. The ret consequently branches to itself and prevents kernel dumps from progressing past dump_savectx(). Tail-call vector_state_store_savectx() so that it returns directly to the original savectx() caller. Reviewed by: br, jrtc27 Approved by: jrtc27 Fixes: https://cgit.freebsd.org/src/commit/?id=d7a393095cfd ("riscv: Vector Extension (RVV) support.") Sponsored by: FreeBSD Foundation Differential Revision: https://reviews.freebsd.org/D59689
The larger CQE exports the RSS header to userspace. Note that this changes the ABI between the driver and library, so the ABI version is bumped. A new user->kernel create_cq request message is defined to disallow a CQE size mismatch between iw_cxgbe and libcxgb4. Sponsored by: Chelsio Communications Co-authored-by: Steve Wise <swise@opengridcomputing.com>
Sponsored by: Chelsio Communications
Sponsored by: Chelsio Communications
This work request permits combining RDMA WRITE and SEND_INV commands into a single work request. Sponsored by: Chelsio Communications
To optimize NVME-oF READ IOPs, use a specialized work request that combines a RDMA WRITE and SEND_INV chain into a single request. Sponsored by: Chelsio Communications Co-authored-by: Potnuri Bharat Teja <bharat@chelsio.com>
- Only insert drain CQEs if a work queue is flushed - Preserve the request opcode of the original request when flushing a request. Use bit 10 of the CQE header word to indicate the CQE is a special drain completion, and save the original WR opcode in the cqe header opcode field. - If a work request chain was posted and needed to be flushed, only the first request in the chain was completed with FLUSHED status. The rest were never completed. - When a CQ is shared by multiple QPs, c4iw_flush_hw_cq() needs to acquire the corresponding QP lock before moving the CQEs into its corresponding SW queue and accessing the SQ contents for completing a WR. - Once a user QP has been flushed, it cannot be flushed again. Sponsored by: Chelsio Communications Co-authored-by: Steve Wise <swise@opengridcomputing.com>
Allocate wait objects dynamically and add a refcount instead of allocating them on the current thread's stack. This permits c4iw_wait_for_reply() to safely fail with an error and mark the device as dead if a reply is not received after C4IW_WR_TO seconds. Once a device is marked dead, future requests fail immediately. Sponsored by: Chelsio Communications
pio_copy() is suppose to be used for write-combining work requests. However, for some architectures like powerpc and arm64, write[ql]() include an explicit memory barrier which will flush the WC buffer and render this path useless. To prevent that use the relaxed variant. Sponsored by: Chelsio Communications
Obtained from: Linux commit f215a3d2448ae77253f0b93dcc37114779f51778
Obtained from: Linux commit 52e124c27e7bfb78980189bdfec049594d7612be Obtained from: Linux commit 2f43129127e62b25f56ff82a37c40b42c0e6f883 Obtained from: Linux commit 7a268a93973c07f93e952d96c2faa88df8ed38d8 Sponsored by: Chelsio Communications Co-authored-by: Potnuri Bharat Teja <bharat@chelsio.com>
Obtained from: Linux commit e00b64f7c54c4cbd88143bbd43e7c3d61a090e5c Obtained from: Linux commit 89944450547334aa6655e0cd4aec8df1897a205a
Obtained from: Linux commit 5f818d676ac455bbc812ffaaf5bf780be5465114
Add await in destroy_qp() so that all references to qp are dereferenced and qp is freed in destroy_qp() itself. This ensures freeing of all QPs before invocation of dealloc_ucontext(), which prevents loss of in use qpids stored in the ucontext. Obtained from: Linux commit f70baa7ee3d1b5a9e66ac7549e31641a656f23c1 Sponsored by: Chelsio Communications
Part of the original Linux commit was already applied during a prior OFED update in commit b633e08c705fe43180567eae26923d6f6f98c8d9, but this portion of the changes to the Linux iw_cxgb4 driver were not included. Obtained from: Linux commit a52c8e2469c30cf7ac453d624aed9c168b23d1af
Obtained from: Linux commit 4c44d4634b5c90993fccca9f155347221df6f877
Obtained from: Linux commit 3840c5b78803b2b6cc1ff820100a74a092c40cbb
These include changes to use newer APIs, cosmetic changes to reduce diffs between the two drivers, and other minor fixes. Sponsored by: Chelsio Communications Co-authored-by: Arjun V <arjun@chelsio.com> Co-authored-by: Krishnamraju Eraparaju <krishna2@chelsio.com> Co-authored-by: Vishal Kulkarni <vishal@chelsio.com>
While here, fix incorrect frees in error paths in c4iw_alloc_context(). Sponsored by: Chelsio Communications
Add an amd64-only, Intel-only diagnostic driver that reads ME Host Firmware Status registers from PCI configuration space without mapping the messaging BAR. Select the HFS register count per device generation, decode sparse HFS1 state, mode, and error fields, and expose raw and summarized status via sysctl. Register an ISA-side hfstsfd probe for supported RCBA/FD2 generations so firmware-hidden HECI functions can be diagnosed. Reviewed by: kbowling, adrian Differential Revision: https://reviews.freebsd.org/D58863
Replace direct IFF_DRV_RUNNING reads in igbv, ixgbe, ixv, iavf, ixl, ice, bnxt Ethernet, aq, enic and the axgbe PCI frontend with iflib_is_running(). Keep each existing mailbox-ready, fault-state, link-state and administrative-up condition: framework admission is not a substitute for device specific readiness. This covers interrupt admission, operational link publication, VF mailbox and VLAN replay, and live-configuration decisions. Remove ifnet temporaries used only to read the driver flags. Replace the em and igc debug dumps' RUNNING/OACTIVE text with one framework admission snapshot. OACTIVE does not mean that the interface or hardware is inactive. Label the result as software admission rather than link state or proof that DMA has stopped, and retain the queue head/tail dumps and separate RS diagnostics. Keep the filter register and requested interface flags in the axgbe PCI promiscuous-mode trace, but remove its unrelated legacy driver flags. Use the parent Ethernet PF context for the bnxt RDMA running check, retaining the link-state check and the existing IB_PORT_ACTIVE policy. Both current RDMA construction paths use the PF netdev, and Ethernet detach synchronously removes the RDMA auxiliary child before iflib frees that context. Declare the direct iflib module dependency. This does not transfer ownership of RDMA queues to iflib. Leave the independent, non-iflib axgbe ARM frontend unchanged. Reviewed by: iflib (gallatin) MFC after: 2 weeks Sponsored by: BBOX.io Differential Revision: https://reviews.freebsd.org/D59601
After a link speed change updates TSO capabilities, request a restart when iflib is running or the interface is administratively up. Leave an administratively down, stopped interface to apply the change at its next initialization. Do not infer initialization from IFF_DRV_OACTIVE, which also remains set after stop and failed initialization. Use iflib_is_running() for software admission and IFF_UP for intent; iflib retains responsibility for quiescing any partially initialized queues before restarting. This also permits a restart request for an administratively up interface before its first initialization attempt. Reviewed by: iflib (gallatin) MFC after: 2 weeks Sponsored by: BBOX.io Differential Revision: https://reviews.freebsd.org/D59599
A typo in the sysctl for RPI0 PWM channel 2 was causing it to write to the wrong register, leaving it misconfigured if used. PR: https://bugs.freebsd.org/bugzilla/show_bug.cgi?id=298301 MFC after: 3 days (problem reported on 14.4) Reviewed by: adrian Reported by: Attila Kover <attila.kover@guardian.co.uk> Differential Revision: https://reviews.freebsd.org/D59644
The ACPI quirk on Framework Laptop 12 firmware was previously clocked at ~620 ms for a Notify 0x80 replayed on power-button press to wake. However, when KMS is not loaded, that very same notification comes in past the 1s we previously allowed. Use the already-defined ACPI_MINIMUM_AWAKETIME (5) seconds as our new boundary so that we can successfully come out of S3 on this hardware. Introduce new tunable hw.acpi.button_replay_window Reviewed by: obiwac, olce Differential Revision: https://reviews.freebsd.org/D59583
[PATCH 04/31] FreeBSD OFED support for DPDK MLX5 PMD a) Expose raw packet capabilities in the core layer to enable a device to report it. Two existing capabilities, scatter FCS and IP CSUM were added to this field for a better user experience by exposing the raw packet caps from one location. This field will serve also for future capabilities for raw packet QP. A new capability was introduced - cvlan stripping, which is the device's ability to remove cvlan tag from an incoming packet and report it in the matching work completion. b) Check device's capabilities and report which raw packet capabilities are supported. c) Allow creating a WQ with cvlan stripping considering device's capabilities. The default value was fixed to disable vlan stripping till was asked explicitly. In addition, allow modification of a WQ to turn on/off this property. d) Enable creating a RAW Ethernet QP with cvlan stripping offload when it's supported by the hardware. e) Add support for creation of a WQ with scatter FCS capability, if this capability is supported by the hardware. Differential revision: https://reviews.freebsd.org/D32178 MFC after: 1 month
Fix is based on upstream Linux commit 2978975ce7f1 ("RDMA/mlx5: Process
create QP flags in one place").
Add IB_QP_CREATE_SCATTER_FCS and IB_QP_CREATE_CVLAN_STRIPPING to the
create_kernel_qp() allowlist, so a kernel RAW_PACKET QP using these flags
is no longer rejected with -EINVAL even though create_qp_common() supports
them. The flags arrived in 358e42ea66e2 and e4cc4fa7cca9, neither of
which updated this mask; upstream only fixed that later, as part of a much
larger refactor that is not worth importing, so the minimal equivalent is
done here.
Sponsored by: NVidia networking
MFC after: 1 month
[PATCH 05/31] FreeBSD OFED support for DPDK MLX5 PMD FreeBSD doesn’t allow userspace to send ioctl call for configuring allmulti mode. Added ALLMULTI flag in mlx5en. In MLX5 PMD, allmulti mode will be enabled/disabled in DPDK only if allmulti is supported by the mlx5en. it won't be changed on network interface side. Differential revision: https://reviews.freebsd.org/D32179 MFC after: 1 month
Import Linux upstream commit 572f46bf947c ("IB/mlx5: Refactor CQE
compression response").
Report the CQE compression capabilities only when the device supports
compression, instead of always advertising supported_format = HASH|CSUM
next to a max_num of 0.
Sponsored by: NVidia networking
MFC after: 1 month
[PATCH 07/31] FreeBSD OFED support for DPDK MLX5 PMD A drop rule is described by an action drop and no destination. If a user specified IB_FLOW_SPEC_ACTION_DROP then set the action to MLX5_FLOW_CONTEXT_ACTION_DROP and clear the destination. Differential revision: https://reviews.freebsd.org/D32181 MFC after: 1 month
Fix is based on upstream Linux commit a22ed86cff36 ("IB/mlx5: Add drop
flow steering rule support").
Initialise is_drop to false so a non-DROP flow cannot inherit stack
garbage, and pass no destination to mlx5_add_flow_rules() for drop rules
instead of one with num_dest = 1, which the flow steering core
dereferences. Both were lost in the original backport.
Derive the destination count from dst as well, so the dont-trap path,
which calls create_flow_rule() with a NULL destination, no longer asks the
core to walk a destination array that is not there.
Sponsored by: NVidia networking
MFC after: 1 month
[PATCH 08/31] FreeBSD OFED support for DPDK MLX5 PMD a) Refactor the netdev notifier registration into a small helper function. This is a pre-step towards having mlx5 IB device over an Ethernet port which doesn't support RoCE. Also, renamed the de-registration helper and the new helper as netdev notifier and not roce, to make it clear this is not only used with roce. b) Rename RoCE related helpers to reflect being Eth ones This is a pre-step towards having mlx5 IB device also over Eth ports where RoCE is not supported. We change the roce enable/disable and roce_lag init/fini function names to have _eth instead of _roce. c) Support RAW Ethernet when RoCE is disabled On some environments, such as certain SRIOV VF configurations, RoCE is not supported for mlx5 Ethernet ports. Currently, the driver will not open IB device on that port. This is problematic, since we do want user-space RAW Ethernet (RAW_PACKET QPs) functionality to remain in place. For that end, enhance the relevant driver flows such that we do create a device instance in that case. Differential revision: https://reviews.freebsd.org/D32182 MFC after: 1 month
Hold a reference on the netdev recorded in dev->roce.netdev and drop it when the entry is replaced or torn down, and register the netdev notifier before the initial interface scan. Otherwise the stored pointer can dangle, and an interface appearing during the scan can be missed. Sponsored by: NVidia networking MFC after: 1 month
PATCH 09/31] FreeBSD OFED support for DPDK MLX5 PMD HyperV is not enabling ROCE by default causing the OFED to not create few SYSCTL variables that publish kernel interface name. Due to this MLX5 PMD is not able to read kernel interface name to get MAC address. Added new sysctl variable in MLX5_EN driver to map the device name to the interface name. Differential revision: https://reviews.freebsd.org/D32183 MFC after: 1 month
Mark the new "ifname" sysctl MPSAFE and use strlcpy() so the interface name it returns is always NUL-terminated. Sponsored by: NVidia networking MFC after: 1 month
[PATCH 10/31] FreeBSD OFED support for DPDK MLX5 PMD Adding UIO driver that exposes the synthetic network device resources to userspace for application such as DPDK can drive it. Added rescind CB capability to the VMBUS driver. Whenever the hypervisor wants to close the device channel, VMBUS on receiving this msg, calls the CB of the device driver (hv_uio in our case). Differential revision: https://reviews.freebsd.org/D32184 MFC after: 1 month
[PATCH 12/31] FreeBSD OFED support for DPDK MLX5 PMD a) When enabling many VFs, the total amount of DMA mappings increase significantly. This causes DMA allocations to take a lot of time since they are serialized in the kernel. As a result the driver enters into fatal condition due to timeout and the system hangs. To recover from this we disable MR cache for VFs. This also resolves Rx-Error MLX5 PMD issue. PFs will still have a full cache and VFs cache can be manipulated as usual after driver load. b) The fast_registration length is used to convey length for memory registrations through UMR which can be of any size up to 2^64. Change the length type to be u64. Differential revision: https://reviews.freebsd.org/D32186 MFC after: 1 month
Import Linux upstream commits 37da2a03c036 ("RDMA/mlx5: Use proper spec
flow label type") and 0f750966dca8 ("IB/mlx5: Add inner spec and IPv6
validation in user's flow attribute list").
The first widens set_flow_label()'s mask and value to u32, so the 20-bit
IPv6 flow label is no longer truncated to 8 bits. The second validates
the ethertype of inner flow specs as well as outer ones, and adds the IPv6
ethertype checking this tree was missing.
Sponsored by: NVidia networking
MFC after: 1 month
Report sw_parsing_caps using the same condition that enables software parsing on the SQ (both eth_net_offloads and swp), so what we advertise matches what the driver actually does. Sponsored by: NVidia networking MFC after: 1 month
Import Linux upstream commit ccc870879027 ("IB/mlx5: Allow creation of a
multi-packet RQ").
Only the create_rq() part was missing: the striding RQ parameters we
validate and store on the WQ never reached the firmware. Program
two_byte_shift_en, single_stride_log_num_of_bytes and
single_wqe_log_num_of_strides, with the two log values encoded relative to
their MLX5_MIN_* bases so the 3-bit firmware fields do not truncate them.
Upstream renamed these WQ context fields later on, so the names used here
are the ones this tree's mlx5_ifc.h still carries.
Sponsored by: NVidia networking
MFC after: 1 month
Drop the unused MLX5_IB_QP_TUNNEL_OFFLOAD flag. Tunnel offload is driven by qp->tunnel_offload_en, so the enum bit was dead and misleading. Also advertise MLX5_RX_HASH_INNER in query_device()'s rss_caps.rx_hash_fields_mask, so user space can discover inner RSS. Sponsored by: NVidia networking MFC after: 1 month
[PATCH 29/31] FreeBSD OFED support for DPDK MLX5 PMD a) Add ib_uverbs_flow_spec_gre to define the rule to match GRE encapsulation protocol. The spec includes the generic specs header, type, size and reserved fields while the filter itself is defined as ib_uverbs_flow_gre_filter and includes: - Checksum present bit, key present bit and version bits in a single 16bit field. - Protocol type field - Indicates the ether protocol type of the encapsulated payload. - Key field - present if key bit is set and contains an application specific key value. b) IB/uverbs: Expose MPLS flow spec to user-kernel ABI header Add ib_uverbs_flow_spec_mpls to define the rule to match MPLS protocol. The spec includes the generic specs header, type, size and reserved fields while the filter itself is defined as ib_uverbs_flow_mpls_filter and includes a single 32bit field named 'label' which consists of: Bits 0:19 - The MPLS label. Bits 20:22 - Traffic class field. Bit 23 - Bottom of stack bit. Bits 24:31 - Time to live (TTL) field. c) IB/mlx5: Add support for GRE flow specification This patch introduces support for the GRE flow spec and allowing the creation of rules based on the protocol and key fields that are part of GRE protocol header. d) IB/mlx5: Expose MPLS related tunneling offloads This patch reports the device's capbilities to offload encapsulated MPLS tunnel protocols to user-space: - Capability to offload MPLS over GRE. - Capability to offload MPLS over UDP. Differential revision: https://reviews.freebsd.org/D32204 MFC after: 1 month
Import Linux upstream commit a93b632c4531 ("IB/mlx5: Fix GRE flow
specification").
Honour the user-supplied GRE protocol mask instead of hard-coding 0xffff,
so wildcard and partial masks work.
Also give flow_spec_data[] 8-byte alignment, so this userspace copy of
struct ib_uverbs_flow_spec_hdr matches the kernel UAPI on 32-bit builds.
The kernel header has always used __aligned_u64 here and only our copy
diverged, so there is no upstream commit for that half. The attribute is
spelled out instead of introducing an __aligned_u64 macro, which
<infiniband/types.h> does not provide and which libbnxtre already defines
for itself.
Sponsored by: NVidia networking
MFC after: 1 month
acpi: Warn if no amdsmu(4) loaded after suspend-to-idle resume If amdsmu(4) is not loaded when entering suspend-to-idle and on an AMD CPU, emit a warning. FreeBSD currently only supports S0ix on AMD CPUs through the SMU. When Intel support is completed, we should check the equivalent for Intel (intelpmc). Reviewed by: olce Sponsored by: The FreeBSD Foundation Differential Revision: https://reviews.freebsd.org/D59672
acpi: Don't compile CPU_VENDOR_{AMD,HYGON} cases on non-x86_64
Fixes build on aarch64.
Fixes: https://cgit.freebsd.org/src/commit/?id=5f68acc931a4 ("acpi: Warn if no amdsmu(4) loaded after suspend-to-idle resume")
Sponsored by: The FreeBSD Foundation
acpi: Fix unused variable error Fixes: https://cgit.freebsd.org/src/commit/?id=9b4caca81de2 ("acpi: Don't compile CPU_VENDOR_{AMD,HYGON} cases on non-x86_64") Sponsored by: The FreeBSD Foundation
acpi: Don't check suspend-to-idle if suspend failed If we e.g. failed to suspend a device and suspend bounced because of that, then we're not expected to have entered a deep sleep state in the first place. In this situation, don't overload the user with irrelevant information. Reviewed by: olce Fixes: https://cgit.freebsd.org/src/commit/?id=5f68acc931a4 ("acpi: Warn if no amdsmu(4) loaded after suspend-to-idle resume") Sponsored by: The FreeBSD Foundation Differential Revision: https://reviews.freebsd.org/D59731
Sponsored by: The FreeBSD Foundation
Use the Hyper-V reset/MAC exchange for 82576 and I350 VFs instead of the native posted mailbox protocol, which the Windows PF does not service. Read the host assigned address through configuration bytes 0x201 through 0x206 only during reset, and use it to identify the matching synthetic hn(4) interface. The operations are local to the VF frontend. Poll hardware link status rather than retaining a native mailbox link handshake. Leave MAC, multicast, promiscuous-mode, and VLAN membership policy with the host. Disable guest VLAN registration and native receive limit requests, and limit the VF to an MTU of 1500 bytes. Preserve accumulated statistics across host resets without counting a counter clear as a wrap. Reject inaccessible register samples and rebase after a reset indication or a disabled transmit queue, including when the PF blocks the queue for malicious driver detection. Document single queue support and host assigned access VLANs. Guest VLAN trunks are not supported: tagged transmissions with the Windows I350 PF driver 14.1.5.0 can disable VF queues even on a trunk-configured port, and the limitation was also reproduced with Windows VF drivers. Such trunks must use the synthetic path with SR-IOV disabled for that virtual adapter. Tested on Windows Server 2025 Hyper-V with both 82576 and I350 PFs. Relnotes: yes Sponsored by: BBOX.io
If a TLS request transmits all but a part of the GMAC at the end of a TLS record, the work request asks the crypto engine to return the calculated GMAC to the driver so it can be sent in a simple TCP packet when the rest of the TLS record is transmitted in the future. However, the offset of the returned GHASH offset was calculated incorrectly in this case causing the driver to not recognize the cached GMAC and instead use a more wasteful work request in the future that encrypted the entire TLS record discarding all but the needed bytes of the trailer. Note that this does not effect correctness, just efficiency. Reviewed by: np Fixes: https://cgit.freebsd.org/src/commit/?id=9e269eafebfc ("cxgbe: Use partial GCM mode for partial TLS records on T7") Sponsored by: Chelsio Communications Differential Revision: https://reviews.freebsd.org/D59711
The Windows PF can disable a VF transmit queue while continuing to report carrier up. Link polling alone then leaves the VF operationally up even though it cannot transmit. The reproduced VLAN failure shows this state with PF driver 14.1.5.0 and an MDD indication in the host trace. Check queue zero from the admin path only while the Hyper-V VF is running with sanitized queues and a completed host handshake. Report operational link down and invalidate the statistics baseline when the queue is disabled. Request recovery through the normal iflib stop/init path only when a fresh, accessible STATUS read reports carrier up. Rate limit requests if the host continues to hold the queue disabled, and leave recovery pending while carrier is down. Document the recovery behavior and clarify why the Hyper-V reset retains the VF-local software reset before its host reset/MAC exchange. Sponsored by: BBOX.io
The Intel shared code can wait for firmware resources while holding its OS abstraction locks. FreeBSD mapped these locks to mutexes, which cannot be held across a voluntary sleep. Concurrent PF rebuilds therefore trigger WITNESS when RSS profile updates contend for the firmware change lock. Map the shared-code lock abstraction to exclusive sx locks. This also covers tunnel and flow-profile operations which can reach the same firmware wait while serialized. Validated with WITNESS on a dual port Intel E835. Sixteen CORE resets rebuilt both PFs without lock warnings, reset failures, or watchdogs. Ten interface down/up cycles and twenty promiscuous-filter cycles also completed cleanly. Reviewed by: erj MFC after: 2 weeks Sponsored by: BBOX.io Differential Revision: https://reviews.freebsd.org/D59339
VLANs configured on the synthetic interface do not otherwise reach the accelerating VF's VLAN callbacks. Those callbacks can be needed for hardware filter membership or for interpreting stripped VLAN tags, even though no vlan interface is attached directly to the VF. Subscribe to VLAN events only in transparent mode and schedule the existing VF task. Snapshot the synthetic interface's VLAN topology under network epoch, then leave epoch before invoking the VF callbacks, which may sleep. Do not acquire hn_lock or configure the VF from a VLAN event handler; the worker applies membership outside the VLAN configuration lock. Keep an applied-VID bitmap under hn_lock and relay only changes. Replay VLANs configured before VF arrival, reconcile changes while acceleration is active, and preserve membership across temporary datapath switches. This relays guest intent; it does not configure host access VLAN policy or overcome PF restrictions on tagged traffic. Block initialization while either interface is detaching. Deregister VLAN handlers before draining work, remove registrations when hn leaves a live VF, and discard the applied bitmap when the VF itself departs. MFC after: 2 weeks Sponsored by: BBOX.io
The transparent VF capability handler ignored the requested change and only copied the VF's enabled capabilities. Forward SIOCSIFCAP to the VF and return its result. Preserve VF capabilities which hn does not expose. Limit advertised capabilities to those supported by the transparent packet path and VLAN relay. Do not inherit VF services such as send tags or the extended capability ioctl when hn has no corresponding methods. At handoff, adopt the VF's enabled offloads without reconfiguring it. Mark the datapath as switching while the VF applies a capability change, since its ioctl may reinitialize the device. If the association is still ready and unchanged afterwards, synchronize hn with the actual VF state even on error and restrict checksum assistance to the forwarded offloads. Republish link state suppressed during the transition when the VF is still ready. Refresh VLAN child capabilities after adoption and when restoring the synthetic path. Run the VF capability ioctl without hn_lock to avoid reversing the VLAN configuration lock order. Hold an ifnet reference, serialize capability changes and VF initialization, and make detach wait with hn_lock released. Revalidate the VF association before publishing the result. Defer VLAN-child capability notifications to the existing VF taskqueue so they also run without hn_lock, and drain that work before ifnet teardown. MFC after: 2 weeks Sponsored by: BBOX.io
Use the key and lookup table lengths returned by GET_VF_RESOURCES when configuring RSS through virtchnl, as DPDK does. The Windows E835 PF advertises a 40-byte key and rejects our fixed 52-byte CONFIG_RSS_KEY request, leaving receive traffic on queue zero. Validate the negotiated lengths before constructing AdminQ messages and publish the lookup table size to iflib. Preserve register-mode RSS selection and its fixed hardware sizes. Use aligned, zero initialized key storage so an RSS kernel's 40-byte key does not leave an uninitialized tail when the PF requests 52 bytes. Validation: normal and RSS enabled iavf module builds passed. On an E835 VF under Hyper-V Server 2025, repeated IPv4 and IPv6 receive tests used all three configured guest RX queues in both transparent hn and non-transparent lagg modes. The RSS key rejection disappeared, IPv4 transmit tests passed, and no TX watchdog fired. Each traffic case used three runs of 16 streams. Obtained from: DPDK (negotiated RSS sizing) MFC after: 2 weeks Sponsored by: BBOX.io
ice_msix_admin() runs as an interrupt filter inside a critical section. ice_rdma_notify_pe_intr() acquires the global RDMA sx and invokes the client event handler, both of which require sleepable thread context. A PE or HMC critical error could therefore panic under WITNESS or sleep from interrupt context. Accumulate OICR causes atomically in the interrupt filter and mark them pending in the driver state. Deliver the notification from the iflib admin task before processing reset events. This preserves the existing ordering, lets an iRDMA-requested reset run in the same admin pass, and coalesces causes from multiple interrupts. MFC after: 2 weeks Sponsored by: BBOX.io Differential Revision: https://reviews.freebsd.org/D59340
ice marked every transmitted packet RS and recorded every last descriptor in its report-status queue. Hardware therefore wrote descriptor status for every packet, and the driver traversed every packet while reclaiming completed descriptors. iflib marks selected packets with IPI_TX_INTR as completion checkpoints. It forces a checkpoint as deferred work or ring pressure grows. Retain EOP on every packet, but set RS and record the descriptor only at those checkpoints. DPDK uses the same sparse-RS design and defaults tx_rs_thresh to 32. Let iflib choose the adaptive interval for FreeBSD. This reduces PCIe and memory traffic while preserving bounded descriptor reclamation. Validated on an E810-XXV in an A-B-A test with five matched four-stream, TSO-disabled transmit runs per phase. Median throughput was 9.413, 9.413, and 9.414 Gbps. Median whole-system CPU was 21.54%, 17.75%, and 21.62%, respectively. The candidate used less system CPU than both exact-baseline phases in every pair. Interrupt rate was unchanged, no transmit watchdog fired, and normal TSO traffic remained line-rate. DTrace confirmed that only iflib-selected packet-final descriptors carried RS after the change. The effect should be more profound at 200Gbps but my DUT is network limited. Reviewed by: gallatin MFC after: 2 weeks Sponsored by: BBOX.io Differential Revision: https://reviews.freebsd.org/D58910
The transmit stop loop cleaned the receive completion ring, and the receive stop loop cleaned the transmit completion ring. Use enic_cq_wq() for transmit queues and enic_cq_rq() for receive queues, matching the queue indices used by the remaining ring bookkeeping. MFC after: 2 weeks Sponsored by: BBOX.io
Do not publish a running interface after vnic_dev_enable_wait() fails. Run the existing stop cleanup for the queues configured before the enable request, then report initialization failure to iflib. Mark those queues as needing cleanup before submitting the enable request so the stopped-state shortcut cannot skip the unwind. MFC after: 2 weeks Sponsored by: BBOX.io
Serialize cached DMA coalescing policy with initialization and schedule its application through the admin task. The old running check preceded if_init() acquiring the context lock, so an intervening down operation could be followed by an unconditional initialization. Use the deferred if-up request so restart permission is checked when the task runs. Changes made while stopped or suspended remain cached for the next initialization. Do not schedule a reset for an unchanged value. MFC after: 2 weeks Sponsored by: BBOX.io
Cache a validated flow control setting without touching hardware when iflib has closed admission. The queues may still be live pending a watchdog stop, or the device may already be stopped or suspended. RX initialization and link setup replay the cached flow control policy. For live updates, decide once under the context lock whether to touch hardware, then mask MDD around SRRCTL writes regardless of a concurrent watchdog closing admission. The context lock excludes actual stop and reinitialization while these writes are in progress. MFC after: 2 weeks Sponsored by: BBOX.io
Raw VSI statistics belong to the firmware-assigned counter index and are reset when firmware recreates the VSI. Reusing a baseline across either event makes a counter decrease look like a full width rollover. Reset the PF and main VSI baselines when rebuilding hardware resources. Start a new baseline when firmware assigns a different PF counter index and whenever it allocates a VF VSI. Ordinary interface reinits which retain the hardware resources continue to retain their statistics. This follows the lifecycle used by ice without placing an unconditional statistics reset in ixl_initialize_vsi(), which also runs during ordinary iflib initialization. MFC after: 2 weeks Sponsored by: BBOX.io Differential Revision: https://reviews.freebsd.org/D59337
Check the result of ixgbe_init_hw() before configuring queues. Report failure through iflib_init_failed() instead of publishing a working datapath after shared-code initialization failed. MFC after: 2 weeks Sponsored by: BBOX.io
Check hardware initialization instead of discarding its result, and preserve the original initialization error across cleanup. Report a failed init to iflib rather than publishing a running datapath. Return normalized FreeBSD errors from attach_post. Reuse the existing native reset cleanup on failure. MFC after: 2 weeks Sponsored by: BBOX.io
Check the receive DMA and frame transfer setup results that were previously discarded. Return EINVAL for invalid ring/writeback addresses rather than reporting success, and stop initialization through iflib_init_failed() when a ring could not be configured. Use the transmit queue count for transmit initialization and the supplied channel index for DMA operations. The driver currently allocates one queue in each direction, so those index/count corrections are latent. MFC after: 2 weeks Sponsored by: BBOX.io
Test the producer and consumer indices after waiting for transmit completion. The post-decrement loop leaves timeout at -1 on exhaustion, so testing timeout == 0 missed the actual timeout and could report one when the final poll completed successfully. Leave the hardware shutdown sequence unchanged. This only corrects the diagnostic; runtime FLR recovery and command-buffer lifetime handling remain separate work. MFC after: 2 weeks Sponsored by: BBOX.io
ice_if_init marks DRIVER_INITIALIZED only after all queue and filter operations succeed, so ice_if_stop intentionally does nothing after an initialization failure. Each failure path must therefore unwind any hardware queues it may have configured before iflib releases their DMA mappings. Tx setup enables firmware scheduler queues one at a time, and Rx enable similarly processes queues incrementally. Route failures from both operations through cleanup paths for both the PF and mirror VSIs. The cleanup helpers tolerate queues which were not configured, so they also cover failures on the first queue. This leaves failed initialization stopped as required by the iflib_init_failed contract. MFC after: 2 weeks Sponsored by: BBOX.io Differential Revision: https://reviews.freebsd.org/D59331
PF and device resets discard all hardware VSI state. ice_rebuild() only recreated the main PF VSI before replaying configuration for every VSI. As a result, replay used stale VF VSI handles. Firmware rejected it with AQ_RC_EACCES and the failure aborted the entire PF rebuild. Re-add each VF VSI before replaying its configuration, and do not publish VFACTIVE until both operations succeed. Clear the initialized state so that the VF must negotiate again. If a VF rebuild fails, leave it inactive and continue so that one guest cannot prevent the PF or sibling VFs from recovering. Mark it initialized only after the resource response is submitted successfully, so the state follows completion of the mailbox handshake. Make that failed state authoritative in the mailbox path. While the firmware VSI is invalid, permit only VERSION and RESET_VF and reject operations which require VSI state. A VFR may complete the hardware reset, but cannot make the PF-owned VSI valid or publish the VF active. Preserve accumulated VF statistics while establishing a new raw hardware sample after reconstruction. This keeps the cumulative totals returned to iavf monotonic across PF resets. Return after rejecting a GET_STATS request for the wrong VSI so that it cannot receive a second success reply. Before a locally initiated reset, notify initialized VFs with RESET_IMPENDING while the mailbox control queue is still alive. Ignore individual send failures so one VF cannot prevent notification of its siblings or the reset itself. Send the event from the common reset preparation path and before directly triggering CORE and GLOBAL resets. Preserve the inactive state across later VF resets. The zero-queue Disable LAN Tx AQ remains mandatory to complete every VFR, but do not clear VFSWR or publish VFACTIVE while the PF-owned VSI remains invalid. Linux ice uses the same separation: generic rebuild excludes VF VSIs, the VF reset path rebuilds them separately, and reset preparation notifies initialized VFs before tearing down the control queues. Validated on an E810-XXV with one and eight host-attached iavf VFs. Repeated PF and CORE resets recovered every VF under traffic without watchdog, MDD, or persistent data-path errors. An additional one-VF test kept traffic active across PF and CORE resets; carrier and traffic recovered automatically and both PF and VF watchdog counters stayed at zero. Two four-queue VFs passed through to a Linux 7.0 iavf guest also recovered carrier and traffic automatically after PF and CORE resets. Simultaneous traffic on both VFs resumed without intervention and the PF watchdog counter remained zero. DPDK 25.11 testpmd, using vfio no-IOMMU and two queues per VF, sustained about 4.55 Mpps per VF before reset. It received reset events for both VFs after PF and CORE resets. The documented ethdev stop, reset, reconfigure, and start sequence restored traffic after each reset. Physical bus mastering remained enabled and the PF watchdog counter remained zero. MFC after: 2 weeks Sponsored by: BBOX.io Differential Revision: https://reviews.freebsd.org/D58905
The VF mailbox handler unconditionally submitted queue-disable commands, including when a VF repeated a request or negotiated after a PF reset had already destroyed its queues. Firmware can reject stale Tx queue metadata, causing a NACK and unnecessary VF recovery. Track queue configuration and enable state across virtchnl operations, and clear it at VF and PF reset. Apply only hardware transitions that are not already complete while preserving progress after a partial failure. Validate queue configurations before mutating hardware so the state maps remain trustworthy. FreeBSD configures Tx hardware in CONFIG_VSI_QUEUES rather than ENABLE_QUEUES, so track Tx configuration separately from Rx configuration and enable state. Linux ice similarly tracks per-VF Tx and Rx queue state and skips redundant transitions. On an E810-XXV, the prior code emitted AQ_RC_EINVAL while configuring a fresh four-queue VF because it disabled never-configured Tx queues. With this change, two VFs completed 20 stop/start cycles each, and recovered across PF and CORE resets with no watchdog, MDD, or AQ errors. MFC after: 2 weeks Sponsored by: BBOX.io Differential Revision: https://reviews.freebsd.org/D58909
The ENABLE_QUEUES handler has already validated the queue selection before attempting to enable Rx hardware. A timeout or unexpected queue state is an operation failure, not an invalid virtchnl parameter. Return VIRTCHNL_STATUS_ERR_ADMIN_QUEUE_ERROR for these failures, matching the disable path. Keep ERR_PARAM for invalid requests and attempts to enable unconfigured queues. Preserve the recorded state of queues which were enabled before a later queue failed. MFC after: 2 weeks Sponsored by: BBOX.io
pagezero() was a plain store loop (bzero), which under the kernel build flags (-mno-vsx -msoft-float) compiles to byte stores. dcbz establishes a zeroed cache line directly in the cache without a read-for-ownership fetch from memory, roughly halving the memory transactions of page zeroing. Measured on POWER9 (Raptor Blackbird, DD2.3, bare metal), zeroing a cold 256 MB buffer with 128-byte scalar loops: byte stores (current libkern memset) 8.6 GB/s doubleword (std) stores 26.8 GB/s dcbz 34.6 GB/s dcbz raises an alignment interrupt on caching-inhibited mappings, and the kernel does not emulate it, so mmu_radix_zero_page() falls back to bzero() for any page whose memattr is not the write-back default. The internal pagezero() callers only touch freshly allocated page-table pages, which are always write-back. Note: dcbz helps only zeroing, where there is no source to read. For page copying it is a pessimization (it adds a redundant zeroing pass on top of the mandatory source read), so mmu_radix_copy_page() is left as a plain bcopy(). Differential Revision: https://reviews.freebsd.org/D59507 Reviewed by: jhibbits
ice_iov_uninit() freed each VF interrupt-map array without returning the reserved indices to the device interrupt resource manager. Repeated VF create and destroy cycles therefore exhausted the PF interrupt map even though no VFs remained. Return the interrupt allocation before freeing its map. Also split software-only VSI release from hardware teardown so failures before ice_initialize_vsi() do not issue invalid RSS, scheduler, and Free VSI commands for an object firmware has never seen. Keep a VF disabled until all of its resources and hardware state have been created successfully. Clear the enabled state before teardown and after any failed add so asynchronous mailbox processing cannot use a partial or freed VSI. Consume VFLR status for inactive VF slots without trying to reset a nonexistent VSI. Track whether firmware currently owns each VSI and clear that ownership after resets. Teardown can then skip AdminQ commands for VSIs which were not rebuilt. Remove every VSI switch filter before firmware teardown, matching Linux and preventing filter-list leaks across create and destroy cycles. This also applies to the PF VSI detach path. Validated on an E810-XXV with two consecutive create and destroy cycles of 128 four-queue VFs. A 16-queue VF could then be created. Two oversized configurations each failed, cleaned back to zero VFs without invalid firmware teardown commands, and were each followed by a successful 16-queue VF creation. MFC after: 2 weeks Sponsored by: BBOX.io Differential Revision: https://reviews.freebsd.org/D58908
VF drivers replay their VLAN filters after a reset and may retry a request whose reply was lost. The PF tracked only a count and sent every requested ID back to the switch. After PF reset replay had already restored the filters, duplicate VID 0 failed with ICE_ERR_ALREADY_EXISTS and NACKed the entire VF batch. Track exact VLAN membership for each VF. Compact requests to unique IDs whose membership changes, enforce the configured limit against those IDs, and update membership after each hardware operation so partial failures cannot undercount filters. Treat already-present adds and already-absent deletes as successful reconciliation and suppress their misleading low-level error dump. Validated on an E810-XXV with a host-attached iavf VF. A three-filter limit was filled with VIDs 0, 1, and 4094. PF and CORE resets replayed all three without a duplicate warning or ADD_VLAN NACK, and DTrace confirmed a three-VID replay reached the PF. A fourth unique VID was rejected without changing the count; deleting an absent VID was a no-op; deleting and replacing a present VID updated the count exactly. A 128 four-queue VF create/destroy cycle also completed cleanly, and a newly recreated VF started with an empty membership map. MFC after: 2 weeks Sponsored by: BBOX.io Differential Revision: https://reviews.freebsd.org/D58909
Intel LPSS I2C controllers enumerated through ACPI rather than PCI never attach on Haswell and Broadwell, so every device behind those buses is lost. On a Dell XPS 13 9343 that hides the I2C HID touchpad and leaves only the PS/2 fallback, which the firmware does not restore after S3. Three causes, all on the ACPI path: Firmware may leave an LPSS function in D3, where its registers read as all-ones and set_controller() fails with "controller error during attach-1". Run _PS0 before mapping them. The PCI path does not need this, which is why the gap went unnoticed. INT33C2, INT33C3, INT3432 and INT3433 are Lynx Point-LP and Wildcat Point-LP, which ig4_pci.c already classifies as IG4_HASWELL; the ACPI path called everything but APMC0D0F an Atom SoC. The functional clock stays gated until bit 0 of IG4_REG_CLK_PARMS is set. Until then the controller accepts writes into the TX FIFO, never drives the bus, raises no interrupts, and every transfer ends in IIC_ETIMEOUT. Linux ungates the same bit in acpi_lpss.c. Doing it in ig4iic_set_config() covers resume as well as attach. With all three in place the touchpad attaches as iichid0/hmt1 with multi-touch and survives suspend and resume. PR: https://bugs.freebsd.org/bugzilla/show_bug.cgi?id=298558 Tested by: Please (XPS 13 9343, Broadwell-U, 15.1-RELEASE-p3) Reviewed by: wulf MFC after: 1 week
Add pci_get_id() and pci_alloc_msi() device methods for aarch64 platforms. Reviewed by: markj (earlier version), andrew Sponsored by: Arm Ltd Differential Revision: https://reviews.freebsd.org/D59169
Add ACPI support to the GICv5 IRS, ITS and IWB drivers. Device parameters are derived from the MADT and IORT tables. Reviewed by: andrew Sponsored by: Arm Ltd Differential Revision: https://reviews.freebsd.org/D59171
CNTPCT_EL0 needs to be read after an ISB in order to ensure a consistent value irrespective of speculative execution. Make existing reads of CNTPCT_EL0 use a new dedicated accessor macro that expands to the correct instruction sequence. Signed-off-by: Kajetan Puchalski <kajetan.puchalski@arm.com> Reviewed by: andrew Sponsored by: Arm Ltd Pull Request: https://github.com/freebsd/freebsd-src/pull/2423
Fix incorrect opcode on CCSIDR_EL1. Signed-off-by: Kajetan Puchalski <kajetan.puchalski@arm.com> Reviewed by: andrew Sponsored by: Arm Ltd Pull Request: https://github.com/freebsd/freebsd-src/pull/2433
Previously, we would be returning AE_OK from acpi_EnterSleepState(). Reviewed by: olce Sponsored by: The FreeBSD Foundation Differential Revision: https://reviews.freebsd.org/D59735
ice: Validate VF virtchnl configuration Virtchnl requests originate in guest-controlled VFs. The existing PF implementation checked message shape but allowed several hardware-facing values through without complete semantic validation. Require the exact RSS key and LUT sizes advertised to the VF, validate every LUT entry, and pass only the validated LUT length to firmware. A short LUT previously caused the AdminQ command to read beyond the received request. Validate queue ring bases, descriptor counts, receive buffer units, maximum frame sizes, and per-VSI consistency before disabling or changing any queue. This also prevents the 32-bit receive ring length from being truncated through a 16-bit validation helper. Advertise the PF frame-size limit in VF resources. Leaving max_mtu zero causes Linux iavf to request a 16382-byte frame, beyond the 9728-byte limit enforced by this driver. Some Linux iavf releases use the advertised maximum as max_pkt_size even when their receive buffers cover fewer bytes within the E810 five-buffer hardware limit. Accept this advisory mismatch for compatibility. ice_setup_rx_ctx() retains the mandatory hardware clamp of RXMAX to five data buffers, as Linux ice does. DPDK sizes its request to the same limit. Preflight the complete interrupt map before changing queue state. Reject invalid ITR indices, traffic on vector zero, duplicate vectors, and queues assigned to multiple vectors. This prevents hostile input from reaching an MPASS or leaving a partially updated mapping. The limits follow the ICE queue context units and the validation policy used by the Intel ICE and FreeBSD ixl PF drivers. MFC after: 2 weeks Sponsored by: BBOX.io Differential Revision: https://reviews.freebsd.org/D59020
ice: Add malformed virtchnl injection points Extend the optional ice(4) failure injection facility with semantic corruption points for queue configuration, RSS keys and tables, and interrupt mappings. Each point mutates an otherwise valid request after the common virtchnl length check. This exercises the PF semantic validators with a real VF while preserving the normal wire format and mailbox path. The queue point selects unaligned Tx or Rx bases, an unaligned or unrepresentable receive buffer, an invalid frame size, duplicate queue IDs, or a bad VSI. The RSS points select short advertised data or an out-of-range LUT entry. The interrupt point selects an invalid ITR, traffic on vector zero, duplicate vectors, or a bad VSI. The points remain absent unless the kernel is built with options DRIVER_FAILPOINTS and retain the existing PF and VF selectors. MFC after: 2 weeks Sponsored by: BBOX.io Differential Revision: https://reviews.freebsd.org/D59021
ice: Enforce VF MAC anti-spoof policy The SR-IOV schema enables MAC anti-spoofing by default, but the driver never programs the VSI security section. A VF can therefore transmit with an arbitrary source address despite the configured policy. Program ICE_AQ_VSI_SEC_FLAG_ENA_MAC_ANTI_SPOOF when the VF VSI is created, and replay the policy when the VSI is rebuilt after a PF or device reset. Fail VF creation or rebuild when firmware cannot install the security policy so an unprotected VF is never published as active. Validated on E810 hardware with host-attached and Linux passthrough VFs. Traffic using the assigned source MAC passed while otherwise identical forged-source frames were dropped. After a PF reset, assigned traffic resumed and zero of ten forged frames reached the peer. An injected MAC anti-spoof update failure left the VF inactive. Destroying and recreating the SR-IOV configuration restored the policy and traffic. MFC after: 2 weeks Sponsored by: BBOX.io Differential Revision: https://reviews.freebsd.org/D59022
ice: Fail closed when VF reset does not complete ice_reset_vf() logs failures to drain PCIe transactions, issue the mandatory zero-queue command, or observe VFR completion, but still publishes VFACTIVE. A VF can resume against reset state which the PF knows is incomplete. Return an error from the reset operation and retain a reset-failed flag when any mandatory stage fails. Reject ordinary virtchnl requests while the failure persists. Publish VFACTIVE only after every stage succeeds. A later VFLR or PF rebuild can recover the VF and clear the failure. Remove tracked queue leaves before clearing their software state. The reset-only AdminQ command drains hardware queues but does not update the shared scheduler database; losing that bookkeeping can strand queue resources across VF teardown and recreation. After a successful VF reset, discard software switch filter state whose hardware rules were reset and clear guest-owned MAC and VLAN tracking. Restore PF-owned anti-spoof policy, the broadcast filter, and the assigned MAC before publishing VFACTIVE. The guest can then replay its own filter state without stale software entries suppressing the firmware requests. Tested on an E810-XXV with a host-attached iavf. Repeated VFLRs rebuilt the base filters with new firmware rule IDs and restored traffic. Forced reset-stage and policy-replay failures remained inactive until recovery, and a PF reset rebuilt the VF and restored live traffic automatically. This follows the conservative reset policy used by ixl(4). MFC after: 2 weeks Sponsored by: BBOX.io Differential Revision: https://reviews.freebsd.org/D59023
Apply upstream changes b9eaf18722221 init_timer() -> setup_timer(), and 86cb30ec07cdc7 setup_timer() -> timer_setup() to get us out of Linux version 4.15 (and earlier) KPIs. Also apply 41cb08555c416 from_timer -> timer_container_of. MFC after: 3 days Reviewed by: jhb, emaste Differential Revision: https://reviews.freebsd.org/D59677
Apply Linux change 55c0fcc3de460 "Convert timers to use timer_setup()" to get us out of Linux version 4.15 (and earlier) KPIs. Also apply 41cb08555c416 from_timer -> timer_container_of. MFC after: 3 days Reviewed by: kib Differential Revision: https://reviews.freebsd.org/D59679
The Linux driver switched away from using this timer in e22979d96a55d. Locally adjust the code to use timer_setup/from_timer in order to get us out of Linux version 4.15 (and earlier) KPIs. Also, right away, apply 41cb08555c416 from_timer -> timer_container_of. MFC after: 3 days Reviewed by: kib Differential Revision: https://reviews.freebsd.org/D59680
The init_timer() was there from day one in (dc7e38ac4da50). It was never needed as the setup_timer() below already did more than just that part of the job. This is part of trying to clear up the Linux pre-4.15 timer KPI from LinuxKPI. MFC after: 3 days Reviewed by: kib Differential Revision: https://reviews.freebsd.org/D59681
ice: Quiesce VFs before device reset Reset preparation notifies cooperative VFs, then immediately releases queue maps and firmware topology. A VF which ignores the notification can continue DMA while the PF tears down the resources which describe it. Assert VFSWR for each configured VF before teardown. Run the mandatory firmware drain serially, disable active receive queues, verify that PCIe transactions have drained, and leave the VF held until its VSI rebuild succeeds. Block ordinary mailbox requests as soon as quiesce begins so a hostile VF cannot re-enable queues in the warning interval. Also clear VFLR status only after VFRD and perform the final Transaction Pending check before publishing VFACTIVE. This follows the VF reset flow in section 4.1.3.3.3 of the Intel E810 Datasheet. Serializing VFs stays below the documented limit of four concurrent VM/VF reset flows. Validated on an E810-XXV with active host VFs and a Linux passthrough VF. Two active VFs recovered together after a PF reset and eight after a CORE reset; every VF returned initialized and traffic resumed. A 128-VF configured passthrough topology also survived a PF reset. With the bhyve process paused while ingress targeted an active Linux VF, the PF completed quiesce, topology teardown, and VSI rebuild without guest mailbox cooperation. The VF remained configured but uninitialized until the guest resumed, then re-handshook and restored traffic. The companion failure-injection commit forced each mandatory drain stage to fail. Every affected VF remained inactive with bus mastering disabled and recovered after SR-IOV recreation. MFC after: 2 weeks Sponsored by: BBOX.io Differential Revision: https://reviews.freebsd.org/D59024
ice: Add VF reset and policy failure injection Extend the optional ICE failure-injection facility with points for the MAC anti-spoof firmware update and each mandatory VF reset stage. The reset points report a failed Tx drain command, VFR timeout, receive queue disable, or final PCIe transaction drain after the corresponding hardware operation. This permits fail-closed state and recovery tests without deliberately leaving live DMA during teardown. The points remain absent unless the kernel is built with options DRIVER_FAILPOINTS and retain the existing PF and VF selectors. MFC after: 2 weeks Sponsored by: BBOX.io Differential Revision: https://reviews.freebsd.org/D59025
ice: Make VF MAC filter requests idempotent VF drivers replay their address filters after reset and may retry a request whose reply was lost. The PF tracked only a count and incremented it after an idempotent hardware add, so duplicate replays eventually exhausted the quota. It then rejected an entire address batch, including the administrator-assigned address. Track exact non-primary MAC filter membership within each VF quota. Validate a complete batch before changing hardware, charge only unique absent addresses, and update ownership after each successful operation. Preserve an administrator-assigned address when the VF is not permitted to change it. Validated on an E810-XXV with a host-attached iavf VF. The configured filter quota was filled, then the complete set was replayed across VFR and PF reset without a duplicate warning or ADD_ETH_ADDR NACK. Deleting an absent address was a no-op. With allow-set-mac disabled, the guest could not remove its administrator-assigned filter, while multicast filter additions continued to succeed. MFC after: 2 weeks Sponsored by: BBOX.io Differential Revision: https://reviews.freebsd.org/D59026
ice: Isolate VFs after malicious-driver detection Consume the per-function MDD latches to attribute transmit and receive events to the offending VF. Treat the global debug registers only as the last-cause diagnostic, add the missing Tx data-protection cause, and select the E830 TCLAN register addresses when required. Block every virtchnl request from an offending VF, reset it, and leave its queues and interrupt mappings unconfigured. Most MDD classes stop a queue, but Tx data protection only drops the offending packet; the reset makes the reported blocked state an actual DMA fence for every class. Complete VFR without restoring resources so a later physical FLR can create a new reset edge and recover the function. Complete VFR before restoring queue and interrupt mappings. E810 does not retain mapping writes while VFSWR remains asserted; retaining the original hardware order prevents an immediate post-attach VFR from leaving queue-map enable clear. E810 also sets a parent PF_MDET latch for an event attributed by a VP_MDET latch to one of its VFs. Do not reinitialize the PF for those event classes. This prevents a hostile VF from flapping PF and sibling traffic while retaining recovery for an unattributed PF queue event. Provide an opt-in auto-reset policy for operators who prefer availability to persistent isolation. Clear reset induced anti-spoof latches before releasing a rebuilt VF. This follows the MDD attribution and reset semantics in section 9.2.2.2.1 of the Intel E810 Datasheet. The per-VF register coverage and E830 TCLAN selection match the Intel ICE driver, while the default fail-closed policy follows ixl(4). Validated on E810 with four four-queue FreeBSD iavf VFs. An invalid tail write blocked only the offender, incremented its Tx MDD counter, caused no PF reinitialization, and left a sibling at 50/50 successful pings. Physical FLR restored the offender and a PF reset rebuilt it without a spurious anti-spoof block. With auto-reset enabled, the offender resumed after the reset while its sibling completed 60/60 pings; all queue-map enable bits remained set after the second VFR. MFC after: 2 weeks Sponsored by: BBOX.io Differential Revision: https://reviews.freebsd.org/D59028
ice: Protect the PF mailbox from flooding VFs Wire the shared code mailbox-overflow detector into the VF lifecycle and virtchnl dispatcher. E830 controllers use their per-VF hardware in-flight-message watermark. On older controllers, attribute a congested mailbox snapshot to its sender, reset it with its queues disabled, and discard its subsequent requests. Advance snapshot accounting even for discarded requests. A physical VFLR, PF reset, or IOV recreation releases the VF. A blocked VF can still submit mailbox messages after reset, so discarding requests does not stop it from replenishing the shared queue. Process at most one initially full mailbox immediately. If producers keep it nonempty, mask only the mailbox interrupt cause and let the periodic admin timer schedule bounded drain work. Keep the shared admin vector enabled so that OICR and other control-queue events can still be serviced. Re-enable the mailbox cause after draining and recheck the queue head for arrivals while the cause was masked. Retry failed reads through the same deferred path instead of treating them as an empty queue. Document both protection models and add failure-injection points for the isolation and scheduling paths. E810-XXV tests with two host-attached iavf VFs injected overflow on VF 0, disabling only that VF while VF 1 completed 60 of 60 pings. A physical VFLR and a PF reset each released VF 0; IOV recreation also reset its cumulative count. Forced persistent mailbox-pending state produced 20 admin and control-queue passes in five seconds, remained timer-paced, and did not interrupt sibling traffic. MFC after: 2 weeks Sponsored by: BBOX.io Differential Revision: https://reviews.freebsd.org/D59029
ice: Report SR-IOV VF status Report the VF MAC, allocated transmit and receive queues, exact trunk VLAN-filter count and capacity, negotiated virtchnl API, configured MAC, VLAN, spoof-check, and promiscuous-mode policy, automatic link-state policy, PF traffic permission, and fault containment through iflib. Expose mirror configuration and active hardware rules, precise malicious-driver isolation and counters, software mailbox-overflow isolation and counters, VF-owned MAC-filter count and limit, and reset diagnostics through a versioned driver.ice extension. Distinguish a failed VF reset from a required VSI rebuild, which may still be pending rather than failed. Keep the namespace schema local to the driver so future extensions need no changes to common network headers or the formatter. Invalidate cached VF handshakes during preparation for an externally initiated device reset, before releasing the context lock to wait for hardware. Mark the VFs as requiring rebuild even if an early PF rebuild failure prevents reaching their VSIs. Software-initiated resets retain their advance notification before invalidating the handshake. The iflib context lock protects VF state and VSI lifetime while constructing the snapshot. The query uses only cached state and does not issue AdminQ commands or read device registers. Validation on E810 with host-attached iavf VFs used ifconfig -v and direct Netlink queries to check handshake, queue, VLAN, policy, MDD, and mailbox state across VF and PF resets. Reviewed by: Pawel Sobczyk <pawel.sobczyk@intel.com>, ziaee Sponsored by: BBOX.io Differential Revision: https://reviews.freebsd.org/D58739
Reviewed by: olce Sponsored by: The FreeBSD Foundation Differential Revision: https://reviews.freebsd.org/D59741
Print out diagnostic information after resuming from suspend-to-idle if we failed to enter S0i3, i.e. the IP blocks that were blocking entry to S0i3. Don't give detailed IP block info for other SMUs than for Phoenix, as I have not had a chance to test these yet and the SMU seems to be very quirky. Reviewed by: olce Sponsored by: The FreeBSD Foundation Differential Revision: https://reviews.freebsd.org/D59673
Reviewed by: olce Sponsored by: The FreeBSD Foundation Differential Revision: https://reviews.freebsd.org/D59674
The VC worker ignores the pending count returned by iavf_process_adminq. If a pass exhausts its message budget, replies can remain in the receive ring without another interrupt to schedule the worker. During init, the iflib admin task cannot provide a backstop because init holds the context lock while waiting for ENABLE_QUEUES to complete. Reschedule the VC worker after a successful pass with messages remaining. Keep the per-pass budget and do not retry AdminQ errors. Stop self-rescheduling once detach clears INITIALIZED; the existing polling and reset-recovery gates continue to exclude asynchronous processing. Validation: reproduced the timeout on a three queue E835 VF under Hyper-V. Temporary tracing showed a successful ENABLE_QUEUES reply already in the guest receive ring after a three-message pass left two messages pending. With the fix and no tracing, a GENERIC kernel with WITNESS and INVARIANTS passed three batches of 30 down/up and MTU 1500/9000 cycles without enable/disable timeouts. MFC after: 2 weeks Sponsored by: BBOX.io
A VF can finish initialization or reset recovery after hn's handoff-time RSS query returned ENXIO. In that case hn suppresses synthetic receive hash metadata, but previously left it disabled even after the VF could answer the queries again. Queue an RSS refresh on a VF link-up notification, using the existing VF worker in both transparent and non-transparent modes. Revalidate the association, active VF path, administrative state, and carrier under hn_lock before querying and reconfiguring synthetic RSS. Coalesce the requests with an atomic flag and retain a request while capability forwarding temporarily excludes the worker. Keep the existing query validity checks and unsupported-query fallback unchanged. This is a one-shot refresh, not a readiness poll. Recovery without a link-up notification, or a query that still fails during the refresh, does not trigger another retry by itself. A later link-up or normal handoff can query again. Consuming the request while the VF path is inactive is intentional: the next handoff queries RSS unconditionally. Continue to reprogram RSS after a successful query even when the cached key and types match. They are updated before host reconfiguration and do not prove that the previous programming succeeded. The event callback only records and schedules work; RSS queries and host reconfiguration remain in sleepable context. No packet-path changes or new timer are needed. With an E835 iavf VF, reproduced an ENABLE_QUEUES timeout and an ENXIO RSS query in both transparent and non-transparent modes. In each mode hn restored mbuf_hash after VF recovery, without another interface reconfiguration. MFC after: 2 weeks Sponsored by: BBOX.io
The SEC QMan interface stores context in context_a and context_b, so add them to the frame queue creation.
The QorIQ Security Engine (SEC) generally fits into the Data Path Acceleration Architecture, accelerating cryptographic operations. Currently this driver does not use the QMan interface, instead relying on the job rings, as Linux also does, to reduce complexity until needed for network protocol acceleration, like IPSec and OpenVPN. Initial testing via `openssl speed -engine devcrpto -evp aes-128-cbc` yields a ~200x throughput improvement, from 105MB/s to more than 20GB/s for 16k block sizes. Differential Revision: https://reviews.freebsd.org/D59581
With the DPAA Security Engine (SEC) KERN_TLS should provide some enhancement over userland software TLS. Also, add IPSEC_SUPPORT, so that ipsec can be loaded if needed.
For Krackan Point, CPU-model-specific matching would not work because amdsmu_match() browses amdsmu_products[] in order and returns the first match, and the Krackan Point's 'struct amdsmu_product' object variant with a 'model' field of 0, indicating that any model matches, is listed before the variant with model 0x70 in amdsmu_products[]. In practice, this means that reporting of IP blocks for Krackan Point model 0x70 only was broken. Specifically, not all the existing blocks were reported and most statistics were not attributed to the right blocks. Fix this by making amdsmu_match() parse amdsmu_products[] in reverse, so CPU-model-generic entries can continue to appear first and new specific ones can be added after them, which is the expected chronological order of additions. While here, since the CPU model is between 0 and 255, change the type used for CPU models to an 'int' and use the special value -1 to skip model match, as there exist CPUs reporting 0 as the model (even if, to our knowledge, only old CPUs seem to be doing that). While here, fix alignement and whitespace in amdsmu_products[]'s initializers. Reviewed by: obiwac Fixes: https://cgit.freebsd.org/src/commit/?id=9c77fb6aaa36 ("amdsmu: Add Krackan Point support") Sponsored by: The FreeBSD Foundation Differential Revision: https://reviews.freebsd.org/D59788
When a graphical application (X.Org display server, Wayland compositor) calls KDSETMODE to set the active VT back to KD_TEXT, vt(4) does not reset the CRTC to point to its framebuffer, leaving the last image of the graphical application visible instead of the text console until the next VT switch. Add a call to vd_postswitch to reset the CRTC, set VDF_INVALID to force a redraw and schedule the flush timer, like vt(4) does in vt_window_switch(). Sponsored by: Defenso Signed-off-by: Quentin Thébault <quentin.thebault@defenso.fr> Reviewed by: vexeduxr Pull request: https://github.com/freebsd/freebsd-src/pull/2308
sound: Standardize mixer_init() call order Device drivers call mixer_init() in a non-standard order - some before pcm_init(), and some after pcm_register(). However, both approaches are potentially risky, and logically weak, since pcm_register() is supposed to be the function that finalizes sound(4) attach. Standardize the call ordering by moving all mixer_init() calls after pcm_init(). This is also necessary for a follow-up patch, which expects pcm_init() to have run first and initialized the PCM lock. Sponsored by: The FreeBSD Foundation MFC after: 1 month Differential Revision: https://reviews.freebsd.org/D59069
sound: Simplify how snd_mixer is fetched and how the cdev is created The primary snd_mixer was reached by accessing the mixer cdev's si_drv1. This is tedious and ugly, so store the mixer in snddev_info->mixer and access it directly. Additionally, create the cdev in a new mixer_make_dev() function (in similar fashion to dsp_make_dev()) in pcm_register(), when everything is initialized, instead of risking potential races because mixer_init() (called before pcm_register()) used to create the cdev. Also add some NULL checks in pcm_register(), to avoid creating a mixer cdev when the driver (e.g., fdt/audio_soc.c) does not create a mixer in the first place, and similarly in pcm_unregister(). Sponsored by: The FreeBSD Foundation MFC after: 1 month Differential Revision: https://reviews.freebsd.org/D59070
sound: Embed mixer cdev in snd_mixer The mixer cdev belongs to the mixer, not to the PCM device, so move it from snddev_info->mixer_dev into a new snd_mixer->cdev field. snd_mixer itself is now included in snddev_info. Also add a MIXER_REGISTERED() macro similar to PCM_REGISTERED(). Sponsored by: The FreeBSD Foundation MFC after: 1 month Differential Revision: https://reviews.freebsd.org/D59071
sound: Defer macio codec volume writes to a task tumbler(4), snapper(4) and onyx(4) write the volume over I2C, and iicbus_transfer() sleeps. This is why mixer_set() drops the mixer lock around MIXER_SET() for non-MPSAFE drivers, relying on Giant to keep them serialized. Store the volume in the softc and let a task do the I2C write with no lock held, so that the mixer method does not sleep at all. The lock dropping for non-MPSAFE drivers will be removed in a follow-up patch. Sponsored by: The FreeBSD Foundation MFC after: 1 month Differential Revision: https://reviews.freebsd.org/D59073
sound: Use snddev_info->lock in place of snd_mixer->lock snd_mixer and snddev_info have a 1:1 relationship. Now that snd_mixer is embedded into snddev_info, it makes even more sense for both to share the PCM lock. The only exceptions to this are MIXER_TYPE_SECONDARY mixers, which still retain a private lock (snd_mixer->priv_lock), because they are attached to the device driver, and not snddev_info. Only snd_emu10kx(4) uses a secondary mixer. A side-effect of this is that the MIXER_SET_LOCK()/MIXER_SET_UNLOCK() mess goes away. These macros were used in the mixer_set*() functions to drop the mixer lock if the driver is Giant-locked and the function can sleep inside MIXER_SET*() methods, and to avoid an LOR before locking PCM to guard channel list traversal. Since mixers now use the PCM lock, drop the channel lock in chn_syncstate() before calling mix_get(), to avoid an LOR. These lines were actually already commented out for years. Sponsored by: The FreeBSD Foundation MFC after: 1 month Differential Revision: https://reviews.freebsd.org/D59074
sound: Retire mixer_hwvol locked variants
Prior to 9a00e0b8ca56 ("snd_uaudio: Do not use snd_mixer->lock as
mixer_lock"), there was a need for mixer_hwvol_mute_locked() and
mixer_hwvol_step_locked(), because the unlocked variants would acquire
the lock, but uaudio_hid_rx_callback() would also hold the lock, so this
was a measure to avoid recursion on snd_mixer->lock. Now that
snd_uaudio(4) has a private mixer lock, the locked variants are not only
unnecessary, but wrong, because we now lock the private lock and not the
snd_mixer one, which is what mixer_hwvol_mute_locked() and
mixer_hwvol_step_locked() expect. Retire the locked variants and call
the regular functions instead.
The unlocked variants take the mixer lock, which is now the PCM lock,
and reach uaudio_mixer_ctl_set(), which takes mixer_lock. Calling them
straight from uaudio_hid_rx_callback() would therefore take mixer_lock
and the PCM lock in the opposite order to the mixer ioctl path, so
record what the HID report asked for and perform the volume change at
the end of the callback, with mixer_lock dropped. The USB stack allows a
callback to drop its transfer mutex (see usbdi.9).
Sponsored by: The FreeBSD Foundation
MFC after: 1 month
Differential Revision: https://reviews.freebsd.org/D59075
sound: Do not set a recording source in mixer_uninit() We currently set the recording source to SOUND_MIXER_MIC during mixer deletion. Apart from the fact that this control might not be present on all devices, it is unnecessary to do that, plus we already set all the volumes to 0 in the mixer_set() call above. Sponsored by: The FreeBSD Foundation MFC after: 1 month Differential Revision: https://reviews.freebsd.org/D59076
sound: Reuse mixer_delete() in mixer_uninit() Sponsored by: The FreeBSD Foundation MFC after: 1 month Differential Revision: https://reviews.freebsd.org/D59077
sound: Improve some mixer return values and their handling Sponsored by: The FreeBSD Foundation MFC after: 1 month Differential Revision: https://reviews.freebsd.org/D59078
sound: Remove unncessary locking in sysctl_hw_snd_hwvol_mixer() The locking around strlcpy() was because of m->hwvol_mixer, but this is just an int, so we don't need to lock in this case. Instead lock only when m->hwvol_mixer is written. While here, add parentheses around the returns. Sponsored by: The FreeBSD Foundation MFC after: 1 month Differential Revision: https://reviews.freebsd.org/D59109
sound: Lock around mixer_set*() in mixer_init() for consistency Sponsored by: The FreeBSD Foundation MFC after: 1 month Differential Revision: https://reviews.freebsd.org/D59110
Treat unsupported RSS queries and RSS_FUNC_NONE as normal reasons to suppress synthetic receive hash metadata. Keep diagnostics for other errors and invalid configurations. Also suppress hash metadata when reconfiguring synthetic RSS fails, since the VF and synthetic settings cannot then be assumed to agree. Correct the hash-query diagnostic name. MFC after: 2 weeks Sponsored by: BBOX.io Differential Revision: https://reviews.freebsd.org/D59691
Use rss_getkey() for PF and VF RSS programming even when options RSS is not configured. The common key is available independently of RSS queue placement and is also used by software hashing. Generating a private key on each initialization defeats that agreement and changes receive hashes across resets. Leave hash-field selection and indirection-table placement unchanged. 82599 and X540 VFs retain their PF-owned RSS configuration. Reviewed by: gallatin MFC after: 2 weeks Sponsored by: BBOX.io Differential Revision: https://reviews.freebsd.org/D59789
Provide the iflib RSS query methods for X550 and newer VFs, whose RSS key and hash types are programmed by the VF. Cache those values and answer queries under the context lock, without register reads or mailbox transactions. Invalidate the snapshot at init, stop, and when an admin check requests recovery after losing the PF mailbox or detecting reset. Reject queries while the snapshot or mailbox is unavailable. Program the common RSS key and cache the values written to the registers. The key remains stable across resets, including live MTU and capability changes for which hn(4) can retain its copy without another VF handoff. Serialize the key in register byte order and report the six supported IPv4/IPv6 TCP/UDP hash selections from the programmed MRQC, rather than claiming that every globally requested hash type is enabled. A single receive queue still performs Toeplitz hashing. Older VFs retain EOPNOTSUPP because their PF-owned RSS settings are not available here. This allows hn(4) to synchronize synthetic RSS with the actual VF settings instead of suppressing otherwise usable receive hash metadata. Reviewed by: gallatin MFC after: 2 weeks Sponsored by: BBOX.io Differential Revision: https://reviews.freebsd.org/D59692
Use rss_getkey() when programming the RSS key through virtchnl regardless of options RSS, matching the register-programming path. The common key is available independently of RSS queue placement; the driver-specific default otherwise makes the chosen key depend on the PF's configuration interface. Remove the unused default-key helper and its declarations. Preserve the negotiated key length, zero padding, hash selections, and indirection-table policy. Reviewed by: gallatin MFC after: 2 weeks Sponsored by: BBOX.io Differential Revision: https://reviews.freebsd.org/D59790
Zero the complete 52-byte RSS key buffer before fetching the common 40-byte key. Both the X722 AdminQ path and the register path consume all 52 bytes; the trailing extended-key bytes must not come from uninitialized stack storage. This matches the zero padding already used by iavf and ice. Remove the unused driver-default key helper and its declaration. Keep the standard key, hash selections, and indirection policy unchanged. Reviewed by: gallatin MFC after: 2 weeks Sponsored by: BBOX.io Differential Revision: https://reviews.freebsd.org/D59791
Use rss_getkey() in the older em multi-queue RSS path, which is reached by 82574 with two receive queues. This replaces a private key generated on each initialization and matches the common-key policy already used by the igb path. Decode each register word with le32dec() and remove the private assembly macro, avoiding signed integer shifts while preserving the register byte order. Leave the redirection table and hash fields unchanged. This does not change single-queue em devices or PF-owned igb VF RSS settings. Reviewed by: erj, gallatin MFC after: 2 weeks Sponsored by: BBOX.io Differential Revision: https://reviews.freebsd.org/D59792
Fetch the common RSS key independently of options RSS instead of generating a private key in non-RSS kernels. Keep the existing software-RSS or round-robin indirection policy and hardware key serialization unchanged. Reviewed by: gallatin MFC after: 2 weeks Sponsored by: BBOX.io Differential Revision: https://reviews.freebsd.org/D59793
Fetch the common key in the PCI iflib frontend regardless of options RSS and make the key API declaration available in both configurations. Retain the existing hash-field selection and queue-placement branches. The independent ARM frontend is unchanged. Reviewed by: gallatin MFC after: 2 weeks Sponsored by: BBOX.io Differential Revision: https://reviews.freebsd.org/D59794
Use the common RSS key instead of generating private random words.
Reverse the bytes before passing the key to the shared code, which copies
the supplied bytes directly into the firmware request. The searcher
consumes this array in reverse order; a raw copy of the common key would
not produce matching Toeplitz hashes.
The hardware ordering is documented by Linux commit
d682d2bdc30650a5c7ce9908ab83ab674b658744 ("bnx2x: byte swap rss_key to
comply to Toeplitz specs"). Keep the conversion in the FreeBSD caller
rather than changing the shared-code interface.
Preserve the config_hash gate, including the PMF-only key programming
on 57710/57711, and leave hash types and indirection policy unchanged.
Reviewed by: gallatin
MFC after: 2 weeks
Sponsored by: BBOX.io
Differential Revision: https://reviews.freebsd.org/D59796
Fill the RSS configuration request with the common 40-byte key instead of fresh random bytes. Preserve the mailbox representation, hash selection, and indirection table. Reviewed by: gallatin MFC after: 2 weeks Sponsored by: BBOX.io Differential Revision: https://reviews.freebsd.org/D59797
Replace the private fixed key with rss_getkey(). Copy the common byte stream directly into the RSS queue-pair context and use the same helper for SIOCGIFRSSKEY, so hardware programming and hn's query agree. Change the helper to fill caller-owned storage instead of returning a pointer to a static key. Check the fixed hardware and ioctl buffer sizes at compile time. Leave RSS masks and queue steering unchanged. MFC after: 2 weeks Sponsored by: BBOX.io Differential Revision: https://reviews.freebsd.org/D59798
Use rss_getkey() for each port's initial key instead of a lazily generated driver-global random key. Remove the obsolete helper and its unsynchronized initialization flag. Make the common-key declaration available without options RSS. The port still retains the key used for hardware programming and the existing SIOCGIFRSSKEY response. Leave hash types and queue placement unchanged. Reviewed by: gallatin MFC after: 2 weeks Sponsored by: BBOX.io Differential Revision: https://reviews.freebsd.org/D59799
Replace the embedded default key with rss_getkey(). Decode each group of four key bytes as a big-endian word; the shared code then serializes those host-order words as little-endian firmware fields. This preserves the existing hardware representation for the common default and allows a configured common key to take effect. Keep the existing multi-queue gate, RSS hash selections, and indirection policy unchanged. Reviewed by: gallatin MFC after: 2 weeks Sponsored by: BBOX.io Differential Revision: https://reviews.freebsd.org/D59800
Replace the embedded key with rss_getkey(), retaining the firmware's reversed 64-bit word-pair order and low-word-first layout within each pair. The common default produces the same mailbox key as the existing constants, while configured common keys no longer get ignored. Leave the RSS enable flags, hash types, and indirection policy unchanged. Reviewed by: gallatin MFC after: 2 weeks Sponsored by: BBOX.io Differential Revision: https://reviews.freebsd.org/D59801
Replace the embedded default with rss_getkey(). Preserve the ICB's little-endian representation of each big-endian key word, then copy the initial 16 bytes of the serialized IPv6 key into the IPv4 key. This retains the old default-key bytes and allows a configured common key to take effect. Leave RSS flags and queue mapping unchanged. Reviewed by: gallatin MFC after: 2 weeks Sponsored by: BBOX.io Differential Revision: https://reviews.freebsd.org/D59802
Replace the embedded key with rss_getkey(). Retain the reversed 64-bit word-pair order and low-word-first layout used by the existing firmware request, matching the qlxgbe conversion. The common default produces the same five firmware words as the old constants. Leave hash selections, RSS enablement, and the indirection mask unchanged. Reviewed by: gallatin MFC after: 2 weeks Sponsored by: BBOX.io Differential Revision: https://reviews.freebsd.org/D59803
Fetch the common Toeplitz key instead of filling it with random bytes during receive-forwarding setup. Check the key size at compile time and retain the existing host-to-network conversion for each hardware word. Leave forwarding, hash selection, and queue mapping unchanged. Reviewed by: gallatin MFC after: 2 weeks Sponsored by: BBOX.io Differential Revision: https://reviews.freebsd.org/D59804
Replace the random register words with rss_getkey() and decode each four-byte chunk with le32dec() before writing the port's key registers. Keep the RSS capability check and existing queue-count and indirection table programming unchanged. Reviewed by: gallatin MFC after: 2 weeks Sponsored by: BBOX.io Differential Revision: https://reviews.freebsd.org/D59805
Replace the repeated hardware-reset key with rss_getkey(). Decode each
eight-byte chunk with be64dec() into the register-word array before
programming the five key registers.
This matches the key serialization in DPDK's nicvf_rss_set_key(), added
by commit 7694fab264ac666d62a4cae04f2979991ce64259
("net/thunderx/base: add RSS and RETA configuration"). A raw copy into
the word array on little-endian hosts would reverse bytes within each
register relative to that convention.
Preserve the CPI-algorithm gate, enabled hash fields, and round-robin
indirection table.
Reviewed by: gallatin
MFC after: 2 weeks
Sponsored by: BBOX.io
Differential Revision: https://reviews.freebsd.org/D59806
Implement SIOCGIFRSSKEY and SIOCGIFRSSHASH for qlnxe and qlnxev using the cached vport RSS configuration. This lets hn(4) synchronize its synthetic RSS configuration with the VF without assuming that every driver programs all of the common RSS hash selections. Serialize the queries with the device lock and reject them while the interface is not running. Convert the firmware request's key words back into network byte order and translate the six representable IPv4/IPv6 TCP/UDP hash selections. Report no hashing when RSS is disabled, including the driver's single-queue configuration. For VFs, require a successful RSS extension response from every hardware function. The overall vport-update reply can succeed while its RSS TLV is missing or rejected; recording that result prevents reporting an unconfirmed configuration. Invalidate the confirmation before each RSS request. Leave existing vport-update return semantics, queue setup, and the receive datapath unchanged. Reviewed by: gallatin MFC after: 2 weeks Sponsored by: BBOX.io Differential Revision: https://reviews.freebsd.org/D59810
Use rss_getkey() for the initial VNIC key instead of generating a per-attachment private key. Preserve the writable rss_key sysctl, its override lifetime, the hardware DMA representation, and all hash and indirection settings. Reviewed by: gallatin MFC after: 2 weeks Sponsored by: BBOX.io Differential Revision: https://reviews.freebsd.org/D59795
Reviewed by: vexeduxr, Quentin Thébault <quentin.thebault@defenso.fr> Sponsored by: The FreeBSD Foundation Differential Revision: https://reviews.freebsd.org/D59814
And replace inline assembly Sponsored by: Arm Ltd Differential Revision: https://reviews.freebsd.org/D59482
arm64/smmu: Move smmu later We need to wait for the interrupt drivers to attach before the smmu driver. As both of these happened at BUS_PASS_INTERRUPT + BUS_PASS_ORDER_MIDDLE there was no guarantee this would happen in the correct order. Fix by moving the smmu driver to BUS_PASS_ORDER_LATE. Sponsored by: Arm Ltd Differential Revision: https://reviews.freebsd.org/D59485
arm64/smmu: Start handling errata Handle the Arm MMU-600 erratum 1076982 where when we poll for a sync command completion we can't use the wait-for-event instruction. This is in preparation for later adding support for polling for sync command completion and is not an issue on stable branches. Sponsored by: Arm Ltd Differential Revision: https://reviews.freebsd.org/D59486
arm64/smmu: Split out waiting for sync completion In preparation for adding a non-MSI method split out waiting for sync completions from smmu_sync. While here fix the loop to wait for sync timeout to the worst case Linux uses. Sponsored by: Arm Ltd Differential Revision: https://reviews.freebsd.org/D59487
arm64/smmu: Support polling for sync completion When the hardware doesn't support MSIs we need to poll to wait for sync operations to complete. Add support for this to support booting on this hardware. Sponsored by: Arm Ltd Differential Revision: https://reviews.freebsd.org/D59488
arm64/smmu: Only MSI poll when cache-coherent Only use the MSI polling method when we support MSIs and the SMMU is cache-coherent. The SMMU writes to a memory location. If it is not cache-coherent then the CPU may read the existing value in its cache and miss the signal the sync operation has completed. Sponsored by: Arm Ltd Differential Revision: https://reviews.freebsd.org/D59489
arm64/smmu: Fix the SMMU-600 errata check Fixes: https://cgit.freebsd.org/src/commit/?id=6dca6e2aa89d ("arm64/smmu: Start handling errata") Sponsored by: Arm Ltd
vmm: Normalize a zero PIT count before starting channel 0 pit_timer_start_cntr0() does not schedule a callout when the initial count is zero. The counter write handler normalizes a programmed zero count only after calling it, leaving an initially unarmed channel 0 without a scheduled timer event. Move the existing normalization before the timer-start call. Retain the historical 0xffff representation of a zero count. Reviewed by: markj MFC after: 2 weeks Sponsored by: The FreeBSD Foundation Differential Revision: https://reviews.freebsd.org/D59749
vmm: Synchronize long-mode state when emulating CR0 writes vmx_emulate_cr0_access() sets EFER.LMA and the IA-32e guest VM-entry control when enabling paging with EFER.LME set, but does not clear them when disabling paging. This can leave an inconsistent guest state that fails VM entry. Update both fields in either direction based on EFER.LME and the CR0 value written to the VMCS. Use the mask-adjusted CR0 value so the resulting state remains consistent with the VMX fixed-bit requirements. Reviewed by: markj MFC after: 2 weeks Sponsored by: The FreeBSD Foundation Differential Revision: https://reviews.freebsd.org/D59756
vmm: Fix a page wiring leak in MOVS emulation When emulating a MOVS from MMIO to guest RAM, the kernel's vm_copy_setup() wires the destination pages. If the subsequent MMIO read fails, emulate_movs() skips vm_copy_teardown(), leaking the page wire references acquired during setup. Run vm_copy_teardown() regardless of the MMIO read result, and only copy the value to guest memory if the read succeeds. Preserve the existing error return. This matches illumos change 13309. Reviewed by: markj Obtained from: illumos 83cd75bb2949d26e6eb38ddefc60fdeed1909643 MFC after: 2 weeks Sponsored by: The FreeBSD Foundation Differential Revision: https://reviews.freebsd.org/D59823
vmm: Re-arm the PIT callout for square wave mode The 8254's square wave mode (mode 3) is periodic, with the same interrupt rate as the rate generator mode (mode 2), but vatpit_callout_handler() only re-arms the channel 0 callout for TIMER_RATEGEN. A guest that programs mode 3 therefore receives a single IRQ0 and no further timer interrupts. Re-arm the callout for TIMER_SQWAVE as well, matching illumos change 13301. Reviewed by: markj Obtained from: illumos 93d78aba5b32996fc2ae893a6237a0d3972f86b2 MFC after: 2 weeks Sponsored by: The FreeBSD Foundation Differential Revision: https://reviews.freebsd.org/D59816
Also explicitly specify .cfi_sections to emit the cfi bytecode into the loadable .eh_frame section. Reviewed by: mchoo Sponsored by: The FreeBSD Foundation MFC after: 1 week Differential revision: https://reviews.freebsd.org/D59838
Reduce very common console messages:
Receive error rdes3=30208000
As observed on the VisionFive v2 hardware after some large transfers.
Differentiate between overflow errors and others. Report the errors when
the length is non-zero (overflow errors).
Also, count errors for netstat purposes.
Reviewed by: mhorne
MFC after: 1 week
Differential Revision: https://reviews.freebsd.org/D59479
It is functional, and there are some RISC-V platforms which can benefit from it. Requested by: Brian Scott <bscott@bunyatech.com.au> Sponsored by: The FreeBSD Foundation
AMD_NPMCS_MAX = 342 (16 core + 6 L3 + 64 DF + 256 UMC). On a Zen 4 EPYC 9654 with 6 core, 6 L3, 16 DF, and 4 UMC counters, only 32 descriptors are needed; the static arrays over-allocate by ~10x. Replace both amd_pmcdesc[AMD_NPMCS_MAX] and per-CPU pc_amdpmcs[AMD_NPMCS_MAX] with mallocarray() sized to the actual registered PMC count: - amd_pmcdesc: allocated in pmc_amd_initialize() - pc_amdpmcs: allocated per-CPU in amd_pcpu_init(), freed in fini() Normalize amd_l3_npmcs and amd_df_npmcs against the AMDID2_PTSCEL2I and AMDID2_PNXC feature bits before computing npmcs_total, so that allocation, registration, and amd_get_msr() row offsets are all derived from the same values. Previously the ternary in npmcs_total excluded L3/DF from the allocation while the globals retained their defaults, causing amd_get_msr() to miscompute DF row offsets when L3 is absent. amd_umc_npmcs comes from CPUID Fn8000_0022h EBX[23:16] (NumUMCCounters) and is zero when the leaf is absent, so no additional feature flag is needed. See AMD64 APM Vol.3 Appendix E. Fix three error-path memory leaks: amd_hwcheck() failure, goto error, and finalize. Reset amd_npmcs = 0 on the error path. Tested on AMD EPYC 9654 (Zen 4, Family 19h Model 11h, 192 threads). Full PMC test suite (IBS/UMCDF/PMC/L3/DF/TSC): 0 failures. Signed-off-by: Osvaldo Janeri Filho <ojanerif@amd.com> Reviewed by: mhorne MFC after: 1 week Sponsored by: AMD Pull Request: https://github.com/freebsd/freebsd-src/pull/2415
s2_tlbi_range under VHE is supposed to clear HCR_EL2.TGE, but it erroneously clears an unrelated bit 27 in TCR_EL2 (HWU61). Since HCR_TGE determines which translation regime will be used by a tlbi, the instruction targets the wrong regime when it is not cleared. Fix the register accessed by the function to actually clear HCR_EL2.TGE. Signed-off-by: Kajetan Puchalski <kajetan.puchalski@arm.com> Reviewed by: andrew Sponsored by: Arm Ltd Differential Revision: https://github.com/freebsd/freebsd-src/pull/2435
This patch adds a driver for the CPU temperature sensor on the jh7110 SoC. The calibration numbers come from the OpenBSD driver but are reworked to produce a result in K rather than C. The temperature is exposed as a sysctl, dev.jh7110_temp.0.temperature but I have also exposed it as dev.cpu.0.temperature because that's where you find it on a RaspberryPi and amdtemp(4), so it's a lot more obvious. (mhorne: Added 'starfive,jh7100-temp' compatible.) Reviewed by: mhorne, bnovkov MFC after: 1 week Differential Revision: https://reviews.freebsd.org/D59478
pmc_ibs_initialize() allocates the ibs_pcpu[] pointer array, and pmc_ibs_finalize() exists to free it, but pmc_ibs_finalize() is never called. Every hwpmc unload on a CPU with IBS therefore leaks one pmc_cpu_max()-sized pointer array. Call pmc_ibs_finalize() from pmc_amd_finalize(), alongside the RAPL, TSC and perf classes. IBS is only initialized on CPUs that support it, so make pmc_ibs_finalize() return early when ibs_pcpu is NULL, making it safe to call when the class was skipped at initialize time, as pmc_rapl_finalize() already is. Tested on an AMD Ryzen 5 5600X (Zen 3, 12 threads) with INVARIANTS. Before the change, each kldload/kldunload cycle leaked one 96-byte M_PMC allocation, and DTrace showed the ibs_pcpu[] allocation from pmc_ibs_initialize() as the only one never freed. After the change, 50 load/unload cycles leave M_PMC InUse and MemUse unchanged, and every allocation made at load is freed at unload. Reviewed by: mhorne Fixes: https://cgit.freebsd.org/src/commit/?id=e51ef8ae490f ("hwpmc: Initial support for AMD IBS") Sponsored by: NLINK (https://nlink.com.br), Recife, Brazil Differential Revision: https://reviews.freebsd.org/D59881
Reviewed by: markj Sponsored by: The FreeBSD Foundation MFC after: 1 week Differential revision: https://reviews.freebsd.org/D59815
Reviewed by: markj Sponsored by: The FreeBSD Foundation MFC after: 1 week Differential revision: https://reviews.freebsd.org/D59815
Reviewed by: markj Sponsored by: The FreeBSD Foundation MFC after: 1 week Differential revision: https://reviews.freebsd.org/D59815
amd64: add md thread flags word Use a hole in struct mdthread. Reviewed by: markj Sponsored by: The FreeBSD Foundation MFC after: 1 week Differential revision: https://reviews.freebsd.org/D59815
amd64: handle #AC in kernel mode Recover from it if PCB_ONFAULT handler is provided. Reviewed by: markj Sponsored by: The FreeBSD Foundation MFC after: 1 week Differential revision: https://reviews.freebsd.org/D59815
amd64: calculate if hardware supports disabling splitlocks Reviewed by: markj Sponsored by: The FreeBSD Foundation MFC after: 1 week Differential revision: https://reviews.freebsd.org/D59815
amd64: support for tracking per-thread 'disable splitlocks' state If hw supports it, on atomic operation that requires exclusive ownership of more than one cache line, #AC is generated. The state is maintained as the arch-private TDF_MD_SPLITLOCK_AC flag. The state is inherited on thread creation from the thread spawning the new one. It is cleared on exec. Reviewed by: markj Sponsored by: The FreeBSD Foundation MFC after: 1 week Differential revision: https://reviews.freebsd.org/D59815
amd64: add userspace control for disabling splitlocks Thread can control it with sysarch(I386_SET_SPLITLOCK). The global default is set with hw.splitlock_force. Reviewed by: markj Sponsored by: The FreeBSD Foundation MFC after: 1 week Differential revision: https://reviews.freebsd.org/D59815
amd64: cache MSR_MEMORY_CTL in pcpu Reviewed by: markj Sponsored by: The FreeBSD Foundation MFC after: 1 week Differential revision: https://reviews.freebsd.org/D59815
amd64: use WRMSRNS immediate form to update splitlock control, when available Reviewed by: markj Sponsored by: The FreeBSD Foundation MFC after: 1 week Differential revision: https://reviews.freebsd.org/D59815
No functional change intended. MFC after: 1 week Sponsored by: The FreeBSD Foundation
They exist only to fill in MODULE_DEPEND() and MODULE_VERSION(), and every consumer passed the same value for all three, so the version range never did anything. Use 1, like the rest of the tree does. MFC after: 1 week Sponsored by: The FreeBSD Foundation Reviewed by: kib, emaste Differential Revision: https://reviews.freebsd.org/D59873
mmu_radix_sync_icache() adds the offset of va within its page to the physical address it gets from mmu_radix_extract_locked(). That address already includes the offset - the extract routines return the physical address of the byte, not of the frame - so the offset is counted twice and __syncicache() is handed frame + 2 * offset. The hash MMU counterpart, moea64_sync_icache(), has to add the offset because PVO_PADDR() yields only the frame. Here the addition is wrong. Fixes: https://cgit.freebsd.org/src/commit/?id=6f0b2a235a13 ("powerpc/pmap: Add pmap_sync_icache() for radix pmap") Reviewed by: jhibbits MFC after: 1 week Differential Revision: https://reviews.freebsd.org/D59870
Turns out that ever since introducing ELFv2 support, it was missing ASLR, it was only used for ELFv1 processes. Reviewed by: jhibbits (via IRC #powerpc64) MFC after: 1 week
The single-vector MSI-X fallback installs the shared interrupt handler, which masks interrupts through INTMS/INTMC around the completion poll. INTx and MSI are unaffected. Only MSI-X track the interrupt mode and skip un/masking Reviewed by: ngie, imp, adrian Differential Revision: https://reviews.freebsd.org/D59637
The Get Log Page request for the error log is clamped to NVME_MAX_AER_LOG_SIZE, but the byte-swap loop iterated ELPE + 1 entries. A controller reporting more than 63 entries makes the loop overrun the 4 KiB log page buffer. Reviewed by: ngie, imp, adrian Differential Revision: https://reviews.freebsd.org/D59626
This fixes a bug where a binary linked using max-page-size=0x200000 can result in a bogus relocation offset when running on a system with a smaller page size. This causes samples to fall outside the image mapping or be translated to the wrong address (resulting in symbol resolution, or incorrect symbol resolution). We noticed this at Netflix because we run a patchset enabling 16k pages on amd64 and have been compiling userspace with a 2MB page size. Since we started doing this profiling userspace binaries has been wonky. Reviewed by: ali_mashtizadeh.com Differential Revision: https://reviews.freebsd.org/D59771 Sponsored by: Netflix
This driver was removed several years ago. Fixes: https://cgit.freebsd.org/src/commit/?id=c1c9764296e5 ("Remove the si(4) driver and sicontrol(8) for Specialix serial cards.")
Fixes: https://cgit.freebsd.org/src/commit/?id=7a323f873662 ("sys: Retire le(4)")
Added nvme_ns_data_format_index() in the nvme, nda, and nvmf host paths as well as nvmecontrol and camdd. Reviewed by: imp, adrian Differential Revision: https://reviews.freebsd.org/D59627
CC.CSS was always writting zero, which is a reserved encoding on a controller that does not support the NVM command set. Select 111b on admin-only controllers and 110b when the I/O command set mechanism is available. Reviewed by: ngie, imp, adrian Differential Revision: https://reviews.freebsd.org/D59628
The receive queue of a PAPR logical LAN is filled in by the hypervisor, so
its fields are big endian, but llan_intr() read the offset and the length of
each frame natively. On a little endian kernel the length of a 134 byte
frame reads as 0x86000000, and ether_input() discards the mbuf because m_len
is not even large enough for an ethernet header. No frame is ever received.
Reproduced on a POWER9 pseries guest with a spapr-vlan interface. Before:
llan0: discard frame w/o leading ethernet header (len -2046820352
pkt len -2046820352)
llan0 1500 <Link#1> 52:54:00:12:34:56 123 118 0 5838733312 7 0
that is 118 input errors out of 123 packets, dhclient(8) never completes and
ping(8) loses every packet, although transmit works because the transmit
path passes the lengths in hcall registers rather than through memory.
Afterwards the interface gets a DHCP lease and ping reports no loss.
MFC after: 1 week
Differential Revision: https://reviews.freebsd.org/D59926
Approved by: jhibbits
Atlantic controllers can retain receive descriptors and their data addresses after their rings are disabled. Reusing or releasing those mappings without invalidating the device cache has caused observed IOMMU and SMMU faults in the referenced Linux reports (7a1bb49461b1, ed4d81c4b3f2 and 7526183cfdbe). Move global cache invalidation out of the per-ring stop routine. Disable every ring first, toggle invalidation once, and wait for its completion indication. Exclude Atlantic A0, as in the upstream workaround. Report a completion timeout rather than silently discarding it. Reviewed by: nprice MFC after: 2 weeks Sponsored by: BBOX.io Differential Revision: https://reviews.freebsd.org/D59852
Stop initialization when hardware setup, ring initialization/start, or datapath start fails. Run the existing best-effort stop/cache/reset cleanup and report the failure through iflib_init_failed(). Do not keep configuring later rings or publish the interface as running. Reviewed by: nprice MFC after: 2 weeks Sponsored by: BBOX.io Differential Revision: https://reviews.freebsd.org/D59853
Fixes: https://cgit.freebsd.org/src/commit/?id=1f38677ba40b ("x86 NOTES: Move shared options from amd/i386 NOTES to x86 NOTES") MFC after: 3 days Sponsored by: The FreeBSD Foundation Differential Revision: https://reviews.freebsd.org/D60011
The int_coal_time and int_coal_threshold handlers differed only in which controller field they updated before reprogramming the feature. Reviewed by: imp, adrian Differential Revision: https://reviews.freebsd.org/D60018
The routine is entered with the namespace lock held and drops it before completing the requests, which preserves the existing locking. Reviewed by: imp, adrian Differential Revision: https://reviews.freebsd.org/D60019
A single handler that takes the counter's offset within struct nvme_qpair in arg2. No functional change Reviewed by: imp, adrian Differential Revision: https://reviews.freebsd.org/D60020
Reviewed by: imp, adrian Differential Revision: https://reviews.freebsd.org/D60021
Catch up with 8452afeb568: now dev.hwpstate_intel.%d.epp accepts values from 0 to 255. Update its description accordingly to sync with the code and the man page.
If userspace provides rate == 0, it causes kernel panic as rate is directly used as a divident, which cannot be zero. Fix it by set the rate to 1 if it is passed as 0. Reviewed by: imp, emaste MFC after: 2 weeks Sponsored by: The FreeBSD Foundation Differential Revision: https://reviews.freebsd.org/D59111
Bytes 102 through 110 of the I/O Command Set Independent Identify Controller data structure were reserved padding. Name them: Boot Partition Capabilities (byte 102), CXL HDM Support Information (byte 103), NVM Subsystem Shutdown Latency (bytes 107:104) and Power Loss Signaling Information (byte 110). Bytes 109:108 stay reserved. CXL HDM Support Information was added in NVM Express Base Specification 2.4, Figure 338; the others appear in 2.3, Figure 328. Report the new fields from nvmecontrol identify. Reviewed by: imp, adrian Differential Revision: https://reviews.freebsd.org/D60000
dpaa2_mc_simplebus_get_devinfo() must not treat the MC resource-container device as a simplebus child. Return NULL when child is sc->rcdev instead of forwarding OFW_BUS_GET_DEVINFO() for that device. MFC after: 2 weeks Differential Revision: https://reviews.freebsd.org/D59836
Stop embedding struct simplebus_softc as the first member of subclass softcs. Use device_get_softc_class() to locate the simplebus portion of the device softc MFC after: 2 months Differential Revision: https://reviews.freebsd.org/D59837
This support was introduced in LLVM 20 and we still support releases with LLVM 19 so it's too early to require LLVM 20. Decay to the non-immediate form when the compiler it too old. This makes the ifunc pointless, but limits the need for other ifdefs. Reviewed by: kib Sponsored by: Innovate UK Differential Revision: https://reviews.freebsd.org/D60023
Creating an executable user-space mapping to write-back memory synchronizes the icache with the page's contents, whether or not those contents have changed since the previous synchronization. Use the pmap private page flag PGA_ICACHE_SYNCED to record that the icache has been synchronized with a managed page's contents and that the page has no writable mappings. When the flag is set, the creation of another executable mapping to the page can skip the synchronization. The flag is cleared when a writable mapping to the page is created, using a single atomic operation that keeps PGA_WRITEABLE and PGA_ICACHE_SYNCED from ever being simultaneously set, and when the page's last mapping is destroyed. Assisted-by: Claude Code (Fable 5.1) Reviewed by: kib, markj Differential Revision: https://reviews.freebsd.org/D59865
Commit 4e0f283fb97a made tpm_tis12_init() wait for TPM_STS_CMD_READY after aborting any command. The wait is implemented by the driver's existing tpm_waitfor_poll() loop, which sleeps with a one-tick tsleep() between status reads. Until now, that loop only ran from the resume and command paths after boot. From tpm_attach() it can panic with "timed sleep before timers are working" when the TPM is attached from ACPI during cold boot and the chip does not report ready on the first status read. Before 4e0f283fb97a, tpm_tis12_init() wrote TPM_STS_CMD_READY and returned without waiting, so the polling loops only ran after boot. tpm_request_locality() had the same latent hazard but its fast path returns before sleeping whenever locality is already active. Nothing calls wakeup() on the channels used by these polling loops, so the sleeps are pure delays. Use pause_sig(), which falls back to DELAY() while the kernel is cold and returns EWOULDBLOCK, a value these loops already tolerate. The c argument to tpm_waitfor_poll() is now unused. It is left in place to keep this change minimal for MFC and can be removed in a follow-up. Reviewed by: kbowling Fixes: https://cgit.freebsd.org/src/commit/?id=4e0f283fb97a ("tpm: Bound TPM 1.2 locality ownership") MFC after: 1 week Sponsored by: The FreeBSD Foundation Differential Revision: https://reviews.freebsd.org/D60050
Patch adding Total Port Shutdown support incorrectly handled a case when this feature was not enabled in the NVM. When TPS bit is not set driver should apply link configuration according to user settings and update the status. Those steps were mistakenly omitted, while the state flag was still set to prevent link renegotation and status update on first attempt to bring interface up with ifconfig. Signed-off-by: Krzysztof Galazka <krzysztof.galazka@intel.com> Reported by: kbowling Reviewed by: kbowling Fixes: https://cgit.freebsd.org/src/commit/?id=0011cd9f8863 ("ice(4): Support Total Port Shutdown on E830 devices") MFC after: 2 weeks Sponsored by: Intel Corporation Differential Revision: https://reviews.freebsd.org/D59578
Correct the byte order at the driver interfaces that do not operate in host order. Encode little-endian Admin Queue VSI and MAC/VLAN fields, and decode firmware-provided queue handles, statistics indices, and event parameters before using them. Decode transmit descriptor writeback before examining it, and construct VLAN descriptor fields in host order for the final descriptor conversion. Store paged HMC descriptors in little-endian form and retain the selected field bits when reading HMC contexts. Keep MMIO values in host order for bus-space accessors and remove the obsolete le16_to_cpu no-op macro. In particular, retain the atomic bus_space_read_8() used for 64-bit registers. On big-endian systems, the unconverted perfect-match filter flag is presented to firmware as 0x0100 instead of 0x0001. Firmware rejects that command with EINVAL, leaving receive traffic functional only in promiscuous mode. Reviewed by: kbowling Approved by: olce (mentor) MFC after: 2 weeks Sponsored by: FreeBSD Foundation, Reliable Computer Systems Lab Pull Request: https://github.com/freebsd/freebsd-src/pull/2408
The radix kremove implementation cleared a kernel PTE without invalidating its TLB entry. Reusing crashdumpmap could therefore keep accessing an old physical page, causing minidumps to contain repeated stale page contents. Invalidate the removed kernel mapping and synchronize newly installed kernel PTEs before they are accessed. Reviewed by: jhibbits Approved by: olce (mentor) Fixes: https://cgit.freebsd.org/src/commit/?id=65bbba25d214 ("powerpc64: Implement Radix MMU for POWER9 CPUs") MFC after: 2 weeks Sponsored by: FreeBSD Foundation Differential Revision: https://reviews.freebsd.org/D59728
libkvm locates the sparse page array after the page-rounded PTE table
size, but the ARM minidump writer emitted only the unrounded size. When
the table size was not page-aligned, libkvm therefore read every dumped
physical page at the wrong offset.
The mismatch was introduced when libkvm began rounding ptesize. It has
affected ARM minidumps since ffdeef323449 ("libkvm: Improve physical
address lookup scaling."). It is exposed when the dumped KVA span is not
a multiple of 4 MiB.
Zero-pad the final PTE page and include the padding in the dump size.
Reviewed by: jhb
Approved by: jhb (mentor)
Fixes: https://cgit.freebsd.org/src/commit/?id=ffdeef323449 ("libkvm: Improve physical address lookup scaling.")
MFC after: 2 weeks
Sponsored by: FreeBSD Foundation
Differential Revision: https://reviews.freebsd.org/D59716
The radix pmap did not implement the minidump pmap callbacks, so radix minidumps were emitted with an empty pmap section. Such dumps do not contain enough information for libkvm to translate kernel virtual addresses. Snapshot the 64KB radix root table in the pmap section, add lower-level page-table pages to the sparse dump, and include pages backing non-DMAP kernel mappings. Keep the bulk direct map out of the dump while retaining the relocated kernel image. PR: https://bugs.freebsd.org/bugzilla/show_bug.cgi?id=298532 Reviewed by: jhb Approved by: jhb (mentor) MFC after: 2 weeks Sponsored by: FreeBSD Foundation Differential Revision: https://reviews.freebsd.org/D59743
Add some additional diagnostic output for IPMI code, particularly on error paths. This has been found to be helpful at $WORK and seems generally useful, so contributing the changes back to upstream. Sponsored by: Dell Technologies Reviewed by: vangyzen@ Differential Revision: https://reviews.freebsd.org/D60091
Service each IBS unit with a valid bit set when fetch and op share an NMI. Otherwise, op samples can be lost. Treat the extra NMI that follows as expected (skip it). Reviewed by: mhorne Fixes: https://cgit.freebsd.org/src/commit/?id=e51ef8ae490f ("hwpmc: Initial support for AMD IBS") Differential Revision: https://reviews.freebsd.org/D60004
hwpmc_amd: add PerfMonV2 global-control path Add support for AMD PerfMonV2 (Family 19h+) core counters, which need both the per-counter EVSEL enable bit and the global GLOBAL_CTL bit set to count. Detects PerfMonV2 at init and switches to v2-specific start/stop/interrupt handlers; older CPUs and L3/DF counters keep using the classic path unchanged. Adds a read-only sysctl, kern.hwpmc.amd_perfmon_v2, to report which path is active. Signed-off-by: Andre Silva <andasilv@amd.com> Reviewed by: Ali Mashtizadeh <ali@mashtizadeh.com> Sponsored by: AMD Differential Revision: https://reviews.freebsd.org/D58256
Some firmware assigns only one bus number to each PCI-PCI bridge. This prevents later SR-IOV VF enumeration when a VF routing ID falls on a bus number already allocated to a sibling bridge. Reserve only the additional bus numbers required by SR-IOV PFs. Enumerate all directly attached functions before child drivers and bridges attach, inspect their device_t objects for SR-IOV, and grow the PCI bus resource through the highest possible VF routing ID. First VF Offset and VF Stride may change when NumVFs changes. Probe every valid NumVFs value and preserve the original setting. When the upstream hierarchy uses ARI, temporarily enable the SR-IOV ARI Hierarchy control in the lowest-numbered PF while sizing, then restore it. Scope active-VF detection to each conventional PCI slot; an ARI bus remains one slot-0 hierarchy. If firmware left VFs enabled on a device, do not modify it and reserve only its active layout. During runtime configuration, consult the PCI bus's owned resource range rather than PCI-PCI bridge registers. This recognizes an existing boot-time reservation beneath both PCI-PCI and host bridges and avoids a second, overlapping bus-number allocation. This avoids consuming bus numbers behind unrelated bridges. The runtime allocation in pci_iov.c remains as a fallback when the boot-time range cannot be enlarged. hw.pci.clear_buses remains useful when firmware assigned a required number to another bridge before enumeration. Validated the targeted implementation on an Intel E810-XXV behind a non-ARI root port. With hw.pci.clear_buses=1 and no global reserve tunable, the PF bridge received buses 1-2 while three unrelated bridges each received one bus. A VF with First VF Offset 0x100 attached as iavf0 at pci0:2:0:0, and detached cleanly. Reviewed by: jhb, manpages (ziaee) MFC after: 2 weeks Sponsored by: BBOX.io Differential Revision: https://reviews.freebsd.org/D58624
Retain the PF-accepted MAC and VLAN settings from the per-port t4iov companion and expose them through the corresponding cxgbe ifnet. Publish, snapshot, and destroy the cache under the existing adapter synchronized-operation mechanism so status queries cannot race IOV configuration or teardown. Track successful t4iov attachment independently of the active VF count. Restrict reporting to the port main VI, return an empty status for a supported but unconfigured PF, and omit status from VF and auxiliary VIs. Reviewed by: jhb Sponsored by: BBOX.io Differential Revision: https://reviews.freebsd.org/D58741
Reviewed by: markj MFC after: 1 week Differential Revision: https://reviews.freebsd.org/D60111
Sponsored by: The FreeBSD Foundation MFC after: 3 days
m_defrag() has pointer-return ownership semantics: success frees the original chain and returns a replacement, while failure returns NULL and leaves the original owned by the caller. rge_encap() compared that pointer as an integer status and could return failure after success, causing rge_tx_task() to free its stale old head. Pass the mbuf by reference, retain the replacement returned by m_defrag(), and publish it to the caller before retrying DMA mapping. A later mapping failure is then cleaned up through the current chain, and a successful transmission uses that same chain for BPF and TX ownership. An unprivileged process can reach the EFBIG branch in mapped-sendfile mode. Signed-off-by: Andrew Griffiths <andrew@calif.io> Reviewed by: adrian, markj MFC after: 1 week Differential Revision: https://reviews.freebsd.org/D60142
Fix build and check return code of intr_ipi_pic_register(). Reviewed by: Sarah Walker <sarah.walker2@arm.com> Differential Revision: https://reviews.freebsd.org/D60071
I have a Logitech mouse that reports 16 buttons. Bump to 32 to be a little more future proof. This increases the size of struct hid_item but this struct is not exported to userland. Reviewed by: wulf Differential Revision: https://reviews.freebsd.org/D59273
Support new versions 4.x and 5.x of Synopsys DesignWare Gigabit Ethernet MAC, often referred to as DWC Quality-of-Service IP Core. In particular: - version 5_30 (0x53) that is found on stm32mp2 arm64 SoC - version 5_40 (0x54) that is found on Spacemit K3 RISC-V RVA23 SoC DWC Ethernet QoS introduced a new scalable multi-channel architecture with up to eight TX/RX channels, enhanced DCB support, TCP segmentation offload (TSO). Thanks to manu@ for start splitting out the DMA code of the older (dwc version 3.x) driver to separate files in 2023. This patch is a continuation of that work to support new core and DMA engine. Thanks to mmel@ for initial review. Initial support covers the same interface capabilities as the version 3 (VLAN_MTU, HWCSUM, HWCSUM_IPV6). Reviewed by: mmel No objection: manu Sponsored by: Innovate UK Differential Revision: https://reviews.freebsd.org/D59191
ARI changes RID interpretation only for the device on a downstream port's secondary bus. pcib_xlate_ari() applied that translation to every config access through an ARI-enabled bridge, including cycles forwarded to a subordinate bus. A non-zero slot on a subordinate bus then panics an INVARIANTS kernel and is misrouted otherwise. Translate only when the access targets this bridge's secondary bus. Reviewed by: kib, jhb Fixes: https://cgit.freebsd.org/src/commit/?id=55d3ea1731d1 ("Add support for PCIe ARI") Sponsored by: AMD Differential Revision: https://reviews.freebsd.org/D60033
The VMX backend did not handle EXIT_REASON_TRIPLE_FAULT, so a hardware triple fault fell through as an unhandled exit: bhyve printed a raw VMX exit dump and aborted, instead of suspending the VM with the reason the software exception path already uses. Add the missing case so the VM suspends with VM_SUSPEND_TRIPLEFAULT and bhyve exits with BHYVE_EXIT_TRIPLEFAULT. This matches illumos change 14664. Obtained from: illumos 83b49c54d9c0766e810b6c8ff849dfb6693fc68a Reviewed by: bnovkov, markj MFC after: 2 weeks Differential Revision: https://reviews.freebsd.org/D59840
acpi_timer: Trim some more leftovers from the ACPI-safe timer The "safe" variant of the hook to read the timer is no longer used and can be removed. Instead, initialize the get_timecount member of acpi_timer_timecounter to the normal hook statically. While here, initialize a few more fields in acpi_timer_timecounter statically. I've kept the name as just "ACPI" instead of "ACPI-fast" now. During device probe there is no longer any reason to alloc the register resource since it is not used, so remove all that. While here, defer registration of the timecounter until attach (kind of odd to do such a thing during probe leaving a window where the timer register was unallocated but in theory could still be read via the timecounter). Reviewed by: cperciva Fixes: https://cgit.freebsd.org/src/commit/?id=00d061855deb ("Garbage-collect ACPI-safe timer and friends") Differential Revision: https://reviews.freebsd.org/D59933
acpi_timer: I/O resource cleanups - Use bus_read_4 and remove explicit bus_space tag and handle - Pass rid by value to bus_alloc_resource_any Differential Revision: https://reviews.freebsd.org/D59934
acpi_timer: Remove unneeded acpi_timer_freq global variable This was just an alias of acpi_timer_timecounter.tc_frequency. While here, register the machdep.acpi_timer_freq sysctl node dynamically only if the driver attaches rather than making the handler fail with EOPNOTSUPP if the driver had not attached. Differential Revision: https://reviews.freebsd.org/D59935
acpi_timer: Add a softc to avoid use of global variables Add a softc and use it to mostly replace the use of global variables in this driver. Simplify the suspend and resume event handlers by saving the old timecounter in the softc and passing the softc pointer to the handlers. Differential Revision: https://reviews.freebsd.org/D59936
The original purpose of this daemon was to route power management requests to userspace, e.g. when the suspend button was pressed, the APM driver posted an event read by apmd(8) that would invoke zzz(8) to suspend. However, on ACPI systems this is handled by devd(8) events instead and running apmd(8) in addition just adds extra complexity. This probably should have been axed when the APM BIOS support was retired in commit 8c576a279ed5. Reviewed by: imp Differential Revision: https://reviews.freebsd.org/D59937
Report failure as if the request had failed. This causes apm(8) to correctly report the resume timer as "unknown" rather than random garbage. Reviewed by: imp Differential Revision: https://reviews.freebsd.org/D59939
This avoids looking it up via a slower path in apmopen. Reviewed by: imp Differential Revision: https://reviews.freebsd.org/D59940
Drop support for queries and commands that are not supported by ACPI's /dev/apm interface. This includes dropping support for enabling/disabling APM BIOS used by /etc/rc.d/apm. Reviewed by: ziaee, imp Differential Revision: https://reviews.freebsd.org/D59941
This simplifies the logic around ACKing suspend requests as there is no longer the potential for multiple listeners, only devd and the acknowledgement via acpiconf -k. Reviewed by: imp Differential Revision: https://reviews.freebsd.org/D59942
cxgbe: Don't query mbuf_nsegs for header-only requests Sponsored by: Chelsio Communications
cxgbe: Various assertions for lengths in KTLS work requests The construction of KTLS work requests is quite fragile, and these assertions ensure that the constructed work requests match the length fields encoded in some of the WR structures. Sponsored by: Chelsio Communications
cxgbe: Don't cache nsegs for KTLS requests KTLS mbufs are initially parsed when they are first enqueued to estimate the number of transmit descriptors needed so that the mbuf is queued until enough descriptors are available. As part of this estimate, the number of DSGL segments required by each KTLS mbuf is calculated. Originally, the count for the first TLS record in a chain was cached in the header mbuf to avoid having to recalculate it when writing out the actual work request for the first TLS record, but this requires duplicating fairly complex logic both when parsing and transmitting requests. Sponsored by: Chelsio Communications
cxgbe: Greatly simplify ktls_wr_len for T7 Don't try to fine-tune the WR size when estimating the work request length when parsing the packet. Use a much simpler worst-case estimate that only depends on a few fields in the mbuf metadata. Sponsored by: Chelsio Communications
When a request needs to drop data from the crypto output (via a split mode request), send the last 16 bytes of input as immediate data instead of via DSGL. Requests with small payloads (16 bytes or fewer) are now sent as immediate data only without any DSGL at all. Sponsored by: Chelsio Communications
Take action in the presence of the GPIO_PIN_PRESET_LOW/HIGH flags. This part of the GPIO interface seems to be unused, but is trivially implemented in our driver. Reviewed by: Brian Scott <bscott@bunyatech.com.au> MFC after: 1 week Sponsored by: The FreeBSD Foundation Differential Revision: https://reviews.freebsd.org/D59991
This provides GPIO programming/configuration at attach time based on the device tree 'pinmux' descriptions. Reference: device-tree/Bindings/pinctrl/starfive,jh7110-sys-pinctrl.yaml Reviewed by: Brian Scott <bscott@bunyatech.com.au> Tested by: Brian Scott <bscott@bunyatech.com.au> MFC after: 1 week Sponsored by: The FreeBSD Foundation Differential Revision: https://reviews.freebsd.org/D59992
Disable these by default as they are less efficient and rarely used. Sponsored by: Chelsio Communications
Network-related commands, library, and kernel.
Update ieee80211_crypto_init_aad() to do what 802.11-2020 says - only mask fc[0] bits 4-6 on data frames, not on management frames. This (with other diffs to actually negotiate MFP and configure ath(4) for MFP + software keys) allows the CCMP path to decrypt CCMP MFP frames in the software path. Differential Revision: https://reviews.freebsd.org/D57799
There is no visible bug fixed as in current tree masks are the same. Fixes: https://cgit.freebsd.org/src/commit/?id=6883b120c53735ff1681ef96d257f376731f56b3
Add const-qualified versions of the NLA iteration helpers to allow walking immutable netlink attribute buffers without discarding const qualifiers. This introduces NLA_NEXT_CONST(), _NLA_END_CONST(), and NLA_FOREACH_CONST() in netlink_snl.h. Signed-off-by: Ishan Agrawal <iagrawal9990@gmail.com> Sponsored-by : Google LLC (GSoC 2026)
Record the Netlink protocol associated with AF_NETLINK sockets when they are created and pass it to libsysdecode during message decoding. Use the protocol to distinguish between Generic Netlink and Route Netlink sockets, ensuring that Generic Netlink decoding is only performed for NETLINK_GENERIC sockets. Signed-off-by: Ishan Agrawal <iagrawal9990@gmail.com> Reviewed by: kp Sponsored-by : Google LLC (GSoC 2026)
Unloading if_ovpn while it's in use by other vnets causes memory leaks and panics. Fix this by reverting VNET_SYSUNINIT and adjusting the SI_SUB initialization order. Reviewed by: markj MFC after: 1 week Differential Revision: https://reviews.freebsd.org/D54175
Revert pf_nl.h part of 017690e50913 and use new libsysdecode build glue that parses enums. Reviewed by: kp, glebius Differential Revision: https://reviews.freebsd.org/D57866
The historical design of sockets is that on a re-connect the disconnect is performed at the socket layer in soconnectat(). Since SMP times this is known to be racy and the function has appopriate comment. I missed that in the recent change. The pr_connect method should normally expect the socket to be already disconnected, however should be able to handle a race where socket is actually connected. Convert the check that incorrectly tried to handle normal path of re-connect into check that handles the race. Reported by: markj Fixes: https://cgit.freebsd.org/src/commit/?id=ece716c5d34728a170f1dfe1b3389c267d6ddd1e
The ```int disable``` parameter is included in the bridge_stop function signature but is not used in the function body. I had noticed this when tracing the driver's path while learning more about the ifnet library. This parameter originally appeared when importing the driver from NetBSD. However, the FreeBSD ifnet library no longer requires an if_stop function. Meaning that the function signature can be changed to only contain needed parameters for our bridge driver. Discussed with: freebsd-net@ mailing list Signed-off-by: Acesp25 <acesp25@freebsd.org> Reviewed by: kp Pull-Request: https://github.com/freebsd/freebsd-src/pull/2290
A mistake from 90ea8e89d9b7 is that in6_pcblookup_internal() was skipped for an inpcb that had unspecified local address. This is incorrect, as such inpcb could have already have a port set, and in_pcb_lport_dest() shall not be called on such inpcb. That could lead to creation of an alised connection in the database. This makes the function almost identical to in_pcbconnect(). While here, fix minor bug of missing INP_ANONPORT. This flag has no use in kernel, but affects netstat(1) output in certain mode. Fixes: https://cgit.freebsd.org/src/commit/?id=90ea8e89d9b751e8b5ae90ef3397883b035788e5 Reviewed by: pouria Differential Revision: https://reviews.freebsd.org/D57987
And make net.inet.ip.portrange.randomized boolean. Reviewed by: pouria, tuexen, markj Differential Revision: https://reviews.freebsd.org/D57291
Enable IFCAP_MEXTPG by default, which may bring performance benefits.
Allow it to be disabled, and when disabled assert that we do not receive
any mbufs with M_EXTPG set. This is useful for testing.
Default the tests to disabling MEXTPG support.
Reviewed by: zlei
Sponsored by: Rubicon Communications, LLC ("Netgate")
Differential Revision: https://reviews.freebsd.org/D58054
The network layer must not pass unmapped (M_EXTPG) mbufs to if_output() of network interfaces without IFCAP_MEXTPG. pf should convert these mbufs by mb_unmapped_to_ext() for such interfaces but it didn't. The problem had occurred on sendfile because sendfile system call uses unmapped mbufs for the file data. Reported by: feld Reviewed by: kp, glebius Differential Revision: https://reviews.freebsd.org/D58021
Some nic drivers (including iflib) do not initialize if_hwassist until after the interface is brought up. If a lagg member is included in a lagg when its not yet been brought up, that will cause lagg to see if_hwassist=0 and will disable all checksum offload, etc, on the interface. This is almost impossible to debug without kgdb or dtrace, as ifconfig does not surface if_hwassist. Fix this by re-calculating lagg caps (including if_hwassist) after adding a port. I encountered this problem when I had a commented-out if_foo1=up entry in my rc.conf that i neglected to uncomment when I was re-configuring a lagg. Sponsored by: Netflix Reviewed by: markj, zlei Differential Revision: https://reviews.freebsd.org/D58062
RFC 4391 (IP over InfiniBand), section 9.3, lays the ND source/target link-layer address option out as type, length (3), two reserved zero octets, then the 20-octet IPoIB link-layer address. The ND code assumed the Ethernet layout (RFC 4861, section 4.6.1) everywhere and read/wrote the address directly after the option header, i.e. two octets early. The option-length sanity check computes 24 for both layouts for a 20-octet address, so the mismatch was silent. PR: https://bugs.freebsd.org/bugzilla/show_bug.cgi?id=296585 Reviewed by: adrian, pouria Sponsored by: VersatusHPC Differential Revision: https://reviews.freebsd.org/D58096
From openssh-portable commit c14709356563. Reported by: cy
Fixes: https://cgit.freebsd.org/src/commit/?id=e7d02be19d40063783d6b8f1ff2bc4c7170fd434 Signed-off-by: Yusuke Ichiki <public@yusuke.pub>
Update nexthop flags with interface link status events and instead of checking link status of interface for every packet only check the reachability flag of the final nexthop. Reviewed by: glebius Discussed with: markj Differential Revision: https://reviews.freebsd.org/D57375
If a nhop gets an interface event, revalidate the nhops and immediately try to recompile existing nexthop groups by replacing unreachable nexthops with reachable ones. If none are available, recompile them back to their normal position in nexthop group slots. Reviewed by: glebius Discussed with: markj Relnotes: yes Differential Revision: https://reviews.freebsd.org/D57389
Provide a name for SCTP sockets. Fixes: https://cgit.freebsd.org/src/commit/?id=8b2b62b49d88 ("sockstat: consolidate unix(4) protocols in the array of protocols")
Fix missing Include which is currently leaked through vnet code. Reported by: bz Fixes: https://cgit.freebsd.org/src/commit/?id=d05d1f256082 ("routing: Subscribe nhops to ifnet link events") Differential Revision: https://reviews.freebsd.org/D57375
Per sys/conf/files this unit is not compiled for a NOIP kernel.
Reported by: Alexander Sideropoulos <Alexander.Sideropoulos@netapp.com> MFC after: 1 week
This is exactly the same as the second part of IPv4's change 136c5e17b61a1/D49153.
Make R-bit per RFC 6275 8.3 and P-bit per RFC 9762 7.1 in Prefix Information option available to userland for future implementations. RFC 9762 7.1: For each interface, the client MUST keep a list of every prefix that was received from a PIO with the P flag set and currently has a non-zero preferred lifetime. Differential Revision: https://reviews.freebsd.org/D56207
- Early return when no new data is delivered - Switching from PRR-CRB to PRR-SSRB only when both SND.UNA advances and no further loss is indicated. - Accounting for sequence ranges SACKed before entering recovery in RecoverFS calculation. - Force a fast retransmit upon entering recovery when prr_out is 0 AND SndCnt is 0. - Set cwnd to ssthresh post recovery. Obtained from: mohnishhemanthkumar_gmail.com Reviewed by: rscheff, tuexen Differential Revision: https://reviews.freebsd.org/D56535 MFC after: 3 months
Add FIB selection logic by introducing ifa_ifwithaddr_fib() to support FIB-specific lookups. Then have ifa_ifwithaddr() wrap it with RT_ALL_FIBS. Also, do the same for ifa_ifwithaddr_check(). Reviewed by: glebius, bnovkov Differential Revision: https://reviews.freebsd.org/D58305
Free mbuf and increase IFCOUNTER_IERRORS if ip_ecn_egress() under geneve_input_inherit() decides to drop the packet. Reported by: Chris Jarrett-Davies of the OpenAI Codex Security Team Reviewed by: markj Differential Revision: https://reviews.freebsd.org/D58361
Sponsored by: NetApp, Inc. MFC after: 1 week Reviewed By: tuexen, #transport, markj Differential Revision: https://reviews.freebsd.org/D58360
When a protocol-specific 'bind_all_fibs' tunable is set to 0, a listening socket will only receive traffic originating from the FIB it was bound to. However, there are no checks to determine whether an address exists in the target FIB when binding the socket, which can lead to a situation where a socket and the address it was bound to belong to different FIBs. Prevent this footgun by looking up the requested address in the current FIB if 'bind_all_fibs' is active and returning an error if the address does not exist. Sponsored by: Stormshield Sponsored by: Klara, Inc. Differential Revision: https://reviews.freebsd.org/D58281 Reviewed by: glebius, pouria, markj MFC after: 2 weeks
In practice, this is not possible, but we are adding it just to be safe. Reported by: markj
pf(4) currently ignores fragment direction (in vs. out)
in pf_frnode_compare() function.
Issue noticed and reported by Frank Denis
OK @bluhm
Obtained from: OpenBSD, sashan <sashan@openbsd.org>, eaa2c80721
Sponsored by: Rubicon Communications, LLC ("Netgate")
Outbound packet which matches rule with source limiter attached,
for example:
source limiter "crash" id 1 entries 10000 limit 1000
pass out from any to any source limiter "crash" keep state
triggers a NULL pointer dereference.
The issue was kindly reported and initial version of fix
submitted by SecBuddyF, Tencent KeenLab.
The submitted diff fixed the issue for failing look up by destination
address in outbound packet. dlg@ also pointed out the change should
be further improved so NULL pointer dereference is avoided when rule
uses nat-to/rdr-to option.
OK dlg@
Obtained from: OpenBSD, sashan <sashan@openbsd.org>, f0f215c11e
Sponsored by: Rubicon Communications, LLC ("Netgate")
The latest version of draft-ietf-tcpm-tcp-ghost-acks changed a condition. This should make no substantial difference, but makei it compliant to the latest version of the specification. Reviewed by: rscheff, Peter Lei MFC after: 3 days Sponsored by: Netflix, Inc. Differential Revision: https://reviews.freebsd.org/D58411
During call to `icmp_verify_redirect_gateway()` ensure using fib-aware source address selection function. Reviewed by: glebius Differential Revision: https://reviews.freebsd.org/D58409
When we drop the prefix lock to call nd6_prefix_offlink() or nd6_prefix_onlink(), make sure to keep the correpsonding prefix structure alive. It is possible for a concurrent nd6_timer() to expire the prefix while the lock is dropped. Reported by: Maik Muench of Secfault Security Reviewed by: pouria, zlei MFC after: 1 week Sponsored by: The FreeBSD Foundation Differential Revision: https://reviews.freebsd.org/D58423
inpcb: declare struct in_conninfo as a single declaration This removes just one level of #define mess that is needed to reach into an inpcbs IPv4 address. And makes the declaration easier to read. No functional change. Reviewed by: markj Differential Revision: https://reviews.freebsd.org/D58273
libdtrace: Fix up translators after struct in_conninfo changes Fixes: https://cgit.freebsd.org/src/commit/?id=698402f4f97c ("inpcb: declare struct in_conninfo as a single declaration")
libdtrace: Fix up one more translator Fixes: https://cgit.freebsd.org/src/commit/?id=d8bcb13b79b4 ("libdtrace: Fix up translators after struct in_conninfo changes")
in_ifprimaryaddr() exists only to support IPv4 multicast usage. Since the adoption of epoch tracking, ifa_ref() is no longer required in its body; that was originally introduced by rwatson in 2009. We could not use __deprecated1() from <sys/cdefs.h> anyway, as IFP_TO_IA() is a macro, not a function. Approved by: glebius (2026-02-26) Reviewed by: adrian, glebius, pouria Differential Revision: D55344
IPv4 multicast currently has the big caveat that it depends on the first assigned IPv4 address on an interface (the so-called "primary address"). in_ifprimaryaddr() only needs to be used by the following: - the 0.0.0.0 booting node input workaround in IGMPv1; - filtering out the node's own reports in IGMPv2; - preserving the source IP where an IGMPv3 report has been looped back; - inferring the default upstream IPv4 interface address for the IP_MULTICAST_IF socket option; - and inferring the source address during ip_output() for a multicast datagram where an interface has been explicitly specified by that option. All of these uses mandate the use of IPv4 source address selection, but FreeBSD does not yet (fully) implement this functionality. Approved by: glebius (2026-02-26) Reviewed by: adrian, glebius, pouria Differential Revision: D55345
If we have source address specified, try to find it by enumerating ifas on specified fib. Reviewed by: glebius Differential Revision: https://reviews.freebsd.org/D58444
Revert UseBlocklist from SSHCFG_ALL to SSHCFG_GLOBAL (with SSHCFG_COPY_NONE), ensuring it can only be set globally in sshd_config rather than within conditional Match blocks, matching historical behavior. Reviewed by: emaste Fixes: https://cgit.freebsd.org/src/commit/?id=bb5c77e9d281 ("OpenSSH: Update to 10.4p1") Differential Revision: https://reviews.freebsd.org/D58520
This may plug minor leaks which no-one has reported. The default IPv6 source
address selection policy list in FreeBSD is usually limited to 9 entries,
and can be readily inspected with ip6addrctl(8). The policy table is
however instantiated for each VNET.
The leak of a pol instance in delete_addrsel_policyent() was already
plugged by @ae in commit-id ecc5c73, so that change has not been merged.
Do not tear down the sxlocks as glebius has requested, and move the
addrsel_policyent{} declarations further up to avoid redundant forward
declarations as glebius requested for stylistic reasons.
Reviewed by: ae, pouria
Sponsored by: Cisco Systems, Inc.
Differential Revision: https://reviews.freebsd.org/D55599
When a TCP timer is stopped, t_timers[] is set to SBT_MAX. Adding the corresponding t_precisions[], if it is not zero, would result in overflows in tcp_timer_next(). To avoid this, skip stopped timers. The problem was identified while debugging uperf by Lukas Book and an initial patch was provided by him. The committed patch was suggested by glebius. The problem can be observed by running netstat -nxptcp and looking for negative timer values and by observing very long running timers in some cases. Reported by: Lukas Book <lkbook@outlook.de> Reviewed by: glebius Differential Revision: https://reviews.freebsd.org/D58484
nd6_ra_input() reads the IPv6 header pointer ip6 before m_pullup(), then uses that pointer afterwards to set nd_ra. When m_pullup() relocates the chain it frees the original first mbuf and returns a new one, leaving ip6 dangling; the subsequent access may be a use-after-free read. The fix writes ip6 from the returned mbuf after m_pullup() inside the conditional if. Reviewed by: pouria Differential Revision: https://reviews.freebsd.org/D58229
nd6_prefix_onlink will enter net_epoch when necessary. Also, exit net_epoch earlier in nd6_prefix_onlink, Because we acquired a reference to ifa, and we got our ifa from the pr->ndpr_ifp, we don't need to stay under epoch. While here, style it. Reviewed by: markj, glebius Discussed with: zlei Differential Revision: https://reviews.freebsd.org/D56129
Add the ability to select source ip address of outgoing packets even when the source ip address is configured on another interface. Also add this new rtnetlink attribute to manual. PR: https://bugs.freebsd.org/bugzilla/show_bug.cgi?id=285422 Reviewed by: glebius, ziaee (manpages) Tested by: ivy, Marek Zarychta <zarychtam@plan-b.pwste.edu.pl> Relnotes: yes Differential Revision: https://reviews.freebsd.org/D58294
Restrict the scope of the struct hc_metrics_lite to the kernel only.
Update the naming to align with other kernel structures and add a tcp_ prefix.
Reviewed by: glebius
MFC after: 2 weeks
Sponsored by: NetApp, Inc.
Differential Revision: https://reviews.freebsd.org/D58440
When configuring the expire timeout to something short, make sure that the
prune time runs at least at that interval. Similarly, when adjusting the
prune interval up, ensure the expire timeout reflect that expected minimum
time also. Finally, restart the callout timer so that the next pruning
happens after the new, expected interval.
Reviewed By: glebius
MFC after: 2 weeks
Sponsored by: NetApp, Inc.
Differential Revision: https://reviews.freebsd.org/D58424
Prior to commit 90ea8e89d9b751e8b5ae90ef3397883b035788e5, this was handled by calling in6_pcbladdr(). Reported by: syzkaller Reviewed by: pouria, glebius Fixes: https://cgit.freebsd.org/src/commit/?id=90ea8e89d9b7 ("netinet6: refactor in6_pcbconnect()") Sponsored by: The FreeBSD Foundation Differential Revision: https://reviews.freebsd.org/D58518
Reported by: Chris Jarrett-Davies of the OpenAI Codex Security Team Reviewed by: pouria, kp Fixes: https://cgit.freebsd.org/src/commit/?id=0361f165f219 ("ipsec: replace SECASVAR mtx by rmlock") MFC after: 1 week Sponsored by: The FreeBSD Foundation Differential Revision: https://reviews.freebsd.org/D58521
Some devices remap the PF queues when entering or leaving SR-IOV. Add opt-in PCI IOV helpers that hold the iflib context lock across the complete stop, driver callback, and restart transaction. Existing drivers continue to use the non-restarting helpers. Sponsored by: BBOX.io
Remove unused rnh_multipath and rib_algo_fixed members. While here, convert rib_dying and rib_algo_init from uint32_t to bool. Reviewed by: glebius Differential Revision: https://reviews.freebsd.org/D58537
Add validation for unused parameter values in the gap between VXLAN_PARAM_WITH_LOCAL_ADDR4 and VXLAN_PARAM_WITH_LOCAL_ADDR6 to prevent panics. PR: https://bugs.freebsd.org/bugzilla/show_bug.cgi?id=297151 Reported by: Robert Morris <rtm@lcs.mit.edu> Reviewed by: markj MFC after: 3 days Differential Revision: https://reviews.freebsd.org/D58552
if_gre(4): Fix races by changing initialization order and locks Treat if_gre like any other network drivers during module initialization by using SI_SUB_PROTO_IF. Also, destroy cloned interfaces via a prison removal callback for gre over udp. PR: https://bugs.freebsd.org/bugzilla/show_bug.cgi?id=275474 Reviewed by: markj Discussed with: glebius Differential Revision: https://reviews.freebsd.org/D57669
if_gre(4): Fix link state announcement in SIOCDIFPHYADDR Since we unlock gre before if_detach() and use slock in gre_clone_modify_nl() there is no need to split if_link_state_change() out of gre_delete_tunnel(). Reported by: markj Fixes: https://cgit.freebsd.org/src/commit/?id=a0d2e5ebaa2e ("if_gre(4): Fix races by changing initialization order and locks")
Migrate to new if_clone KPI and implement netlink support for gif(4). Also break GIFSOPTS ioctl logic out of gif_ioctl. Reviewed by: markj Differential Revision: https://reviews.freebsd.org/D57666
A RST segment can be sent in response to (a) received segment or (b) by the upper layer protocol. The SEG.SEQ validation consists of two checks: (1) the in-window check of SEG.SEQ and (2) the exact match check of SEG.SEQ. For the in-window check (1), the left edge of the window needs to be based on tp->last_ack_sent to cover the delayed ACK case, whereas the right edge needs to be based on tp->rcv_nxt + tp->rcv_wnd. This both assumes that tp->rcv_wnd is not zero. For the special case of tp->rcv_wnd being zero, add checks against tp->last_ack_sent for (a) and on tp->rcv_nxt for (b). This applies to all TCP stacks. When the exact match (2) of SEG.SEQ is performed, it should be based on tp->last_ack_sent for (a) and on tp->rcv_nxt for (b). To cover both, check for both. Add this only to the base stack, since the RACK and BBR stacks already do this. PR: https://bugs.freebsd.org/bugzilla/show_bug.cgi?id=296594 Reviewed by: rscheff MFC after: 3 days MFC to: stable/14 MFC to: stable/15 Sponsored by: Netflix, Inc. Differential Revision: https://reviews.freebsd.org/D58594
This change is intended to address @glebius comments from the original D55663.
ip6_hdr_pseudo{} is referenced by certain OpenBSD OCF related components. I am
using __aligned(4) and not __packed as urged by the late Hans-Petter Selasky.
Use C99 types and style. We must eat the churn now cross-BSD compatibility
is "Fade to Black".
Put _Static_assert under #ifdef INVARIANTS to not disrupt regular compilation,
as this resides in a commonly included header file.
This brings xform_tcp.c into line with possible future OCF related imports.
Add support for allowing IPv4 multicast groups to be joined on IPv6 sockets, as a number of applications began to rely on this over the years, despite it only ever having been a convenience which appeared in Solaris & Linux over the course of the 00s decade. It is limited to any-source joins (ASM). To avoid further quibbling over the meaning of the term "undocumented" as it applies to this change, I have chosen to use the wording "non-IETF-ratified extension" in comments, with reference to the updated ip6(4) man page. PR: https://bugs.freebsd.org/bugzilla/show_bug.cgi?id=https://bugs.freebsd.org/bugzilla/show_bug.cgi?id=193246
Comment updated only. No functional change. It is unrealistic to expect that this feature will ever be resurrected from the legacy KAME tree, given historical divergence, and that applications which really need to consume all group state (e.g. proxies) will either join on a per-group basis, or use link-layer mechanisms anyway. It was also very poorly documented to begin with.
1. EFAULT was happening because sooptcopyin() from inp_join_group() was seeing the user-space thread descriptor in the faked-up sockopt. So, do not attempt a user copyin(); defer to C99 initialization nulling sopt_td for us to force a KVA memcpy(). 2. It seems necessary to byte-swap ipv6mr_multiaddr.s6_addr32[3] on amd64 for similar reasons as to how the user-space initialization needed for passing an IPv4-mapped group address also requires byte-swapping of the 0x0000FFFF field for s6_addr32[2]; it is a direct assignment to a integer member of a struct, NOT a memcpy(). 3. The assignment to imr_interface within in6_v6_mreq_to_v4() was obfuscated by a cast back to its own type due to use of the IA_SIN() macro. Elided. With this change, the feature gap seems to be closed; tested with a simple link-scope IPv4 group under 224.0.0.0/24 with an mlx5(4) SR-IOV VF in bhyve. PR: https://bugs.freebsd.org/bugzilla/show_bug.cgi?id=https://bugs.freebsd.org/bugzilla/show_bug.cgi?id=193246 Differential Revision: https://reviews.freebsd.org/D58590
cmd_securelevel is the securelevel at which the call should be denied. pf (write) calls should be denied at level 3 or up (not at 2 or up as it was), so increment these all by one. PR: https://bugs.freebsd.org/bugzilla/show_bug.cgi?id=296838 MFC after: 4 weeks Sponsored by: Rubicon Communications, LLC ("Netgate") Differential Revision: https://reviews.freebsd.org/D58377
pf assumes that network groups and network interfaces share a namespace (that is, a name is unused, a group or an interface, never both a the same time). Unfortunately this assumption was broken when interface renaming was introduced. Attempt to cope with this rather than panicking. Note that this is a band-aid, not a full solution. The correct fix is for the network stack to go back to enforcing a single namespace for groups and interfaces. PR: https://bugs.freebsd.org/bugzilla/show_bug.cgi?id=297220 Reported by: Robert Morris MFC after: 1 week Sponsored by: Rubicon Communications, LLC ("Netgate")
This fixes a bug where we do not report all lanes when a NIC configures a breakout. Eg, we reported all 4 lanes when a NIC configured the optics as 1x400g, but only printed the first lane's strength when configured as 4x100g. Fix this by actually parsing the active lane count, rather than pulling it from the default descriptor. While here, optionally print page 10h when -vvvv is specified. This aids in determining how a breakout is configured. I put it under an extra level of verbosity, as I don't want to let things get out of hand printing CMIS pages. Sponsored by: Netflix Reviewed by: kib, sumit.saxena_broadcom.com Differential Revision: https://reviews.freebsd.org/D58263
The SIOCGETSGCNT handler may be invoked in this scenario, and if no router has initialized the lookup table, we'll have mfct->mfchashtbl == NULL. PR: https://bugs.freebsd.org/bugzilla/show_bug.cgi?id=297148 Reported by: Robert Morris MFC after: 1 week Sponsored by: The FreeBSD Foundation
This should assist setups that connect a lot of nodes to ipfw: and then distribute traffic with ipfw(4) tablearg feature. Reviewed by: pouria Differential Revision: https://reviews.freebsd.org/D58547
iflib: restore TX watchdog functionality Since f6afed726b00 the TX-hang check in iflib_timer() has required a queue state other than IFLIB_QUEUE_IDLE, but nothing ever sets IFLIB_QUEUE_WORKING, so IFLIB_QUEUE_HUNG has been unreachable ever since: stalled TX queues are not detected, not reported, and not reset - the TX watchdog of every iflib(4) driver has been dead code. Instead of resurrecting the queue-state machine, detect the hang directly. A transmit queue is frozen while it holds descriptors the hardware has not reported as completed and none were reclaimed over a timer period. Being frozen is not a fault: the hardware may defer marking descriptors as completed indefinitely. The check therefore arms only when a frozen queue also takes on new work, while the link is up, no pause frames were received and no doorbell is pending; and it acts only after the queue has stayed frozen for net.iflib.tx_watchdog_periods consecutive periods. It then asks the hardware through the driver's read-only credits peek (isc_txd_credits_update with clear=false, the same call the mp_ring can_drain callback makes routinely): if completions are ready but were not harvested for this long, the completion interrupt went missing - kick the queue's task instead of resetting; if the hardware reports nothing although the queue kept receiving work, it is hung and the existing watchdog reset machinery takes over. Neither software counters alone nor mere persistence of unharvested work can make this decision. iflib reclaims lazily (up to isc_tx_nsegments completed descriptors stay unharvested indefinitely) and defers report-status requests, so "descriptors in use" and "no cleaning progress" are normal states of an idle healthy queue. And hardware that coalesces completion reports (e.g. 8254x, TXDCTL.WTHRESH) legitimately withholds the last one of a quiet queue indefinitely, so a zero credits peek is a normal idle state, not a hang indicator: arming on persistence alone reset healthy interfaces on every traffic lull (field-tested on 82541PI). Only growth across frozen periods separates a wedged queue from a coalescing one. The threshold is a threshold in time, not in device work: a period is one iflib_timer interval (hz/2 by default), so at the default of four periods the verdict falls after roughly two seconds. It was calibrated from counter traces on that old and slow hardware, where healthy coalescing always cleared within two periods; newer hardware reports completions far sooner and leaves the frozen state earlier, so the default needs no recalibration for more modern devices. Setting the sysctl to zero disables the check. A queue whose link is down is never flagged - preserving what f6afed726b00 fixed. The new per-queue state goes into padding the transmit queue structure already had, rather than next to the counters it is derived from: that region is packed, so an insertion there would grow the structure. What is left of that padding is now spelled out instead of being implicit. The size of the structure is unchanged on amd64, arm64, riscv64, i386 and armv7. The IFLIB_QUEUE_* states no longer participate in the watchdog decision; they will be removed in a followup commit. PR: https://bugs.freebsd.org/bugzilla/show_bug.cgi?id=220997, https://bugs.freebsd.org/bugzilla/show_bug.cgi?id=239240 Fixes: https://cgit.freebsd.org/src/commit/?id=f6afed726b00 ("iflib: Prevent watchdog from resetting idle queues") Suggested by: gallatin (mxge-style detection) Reviewed by: adrian, markj MFC after: 1 month Differential Revision: https://reviews.freebsd.org/D58266 Assisted-by: Claude Code (Fable 5, Opus 5)
iflib: remove the unused TX queue state machine The previous commit stopped using ift_qstatus and the IFLIB_QUEUE_* states for the TX watchdog decision, leaving only dead stores. Remove the field, the states, and all assignments. The byte the field frees stays behind as explicit padding. No functional change. Reviewed by: gallatin, markj MFC after: 1 month Differential Revision: https://reviews.freebsd.org/D58282 Assisted-by: Claude Code (Fable 5, Opus 5)
iflib: clear the deferred TX descriptor state when a queue is stopped Stopping an interface frees the queued mbufs and zeroes a transmit queue's descriptor accounting, but the three counters that track descriptors deferred to a later doorbell write or report-status request are not cleared there: they only reach zero when the code that acts on them runs. After a reset they therefore describe descriptors that no longer exist, until enough new traffic flushes them. The consequences are small - one doorbell written from a stale count, and a report-status request on the first packet after the reset - but the state is simply wrong, and the transmit-hang check in iflib_timer() reads one of them. MFC after: 1 week Assisted-by: Claude Code (Opus 5)
iflib: Track queue datapath lifecycle Track whether iflib queue mappings may still be accessed by the device. Keep the state private to iflib and conservative: an unknown or failed device must pass through IFDI_STOP() before mappings are reused or released, while a device known to be stopped need not receive another hardware stop. Enter the starting state before IFDI_INIT(), publish running only after receive buffers and framework state are ready, and stop hardware if receive-buffer setup fails after driver initialization. Do not initialize an administratively-down interface merely because its MTU, capabilities, VLAN configuration, or media changed. Preserve successful retries for an administratively-up interface whose previous initialization failed. This state describes ownership of iflib datapath mappings only. It deliberately makes no claim about firmware queues, administrative DMA, PCI power state, or whether a driver can safely elide a hardware reset. Reviewed by: iflib (gallatin) MFC after: 2 weeks Sponsored by: BBOX.io Differential Revision: https://reviews.freebsd.org/D59328
iflib: Own queue quiescence during power transitions Perform a terminal datapath stop before suspend and shutdown callbacks, then drain the private configuration taskqueue before entering low power. Track power state independently of queue ownership and prevent built-in admin, IOV, LED, and media-status callbacks from accessing a suspended device. Restore driver-specific state while the datapath remains stopped. Initialize it exactly once on resume when the interface is administratively up, and keep an administratively-down interface stopped. Roll back the driver when suspend or child suspension fails. Add ifdi_power_prepare() for policy which must be established before the terminal stop. Use it to snapshot ixgbe(4) wake policy and preserve X550EM PHY ordering, and remove the duplicate stop from aq(4). Reviewed by: iflib (gallatin) MFC after: 2 weeks Sponsored by: BBOX.io Differential Revision: https://reviews.freebsd.org/D59329
iflib: Reject media changes during suspend iflib gates its built-in administrative and media-status callbacks once a power transition starts, but iflib_media_change() could still invoke a driver while the device was suspending or suspended. Several drivers perform PHY or firmware I/O directly from this callback. Return EBUSY before invoking IFDI_MEDIA_CHANGE() unless the device is active. ifmedia then restores the prior selection, avoiding both suspended hardware access and an unvalidated configuration that would need to be replayed during resume. Validated with device suspend on 82579LM, I210, and I225-IT controllers. Media-selection requests returned EBUSY on every suspended device. Resume restored the linked management interfaces at 1 Gbps with working traffic and no watchdogs; unconfigured interfaces retained their prior admin and link state. Reviewed by: iflib (gallatin) MFC after: 2 weeks Sponsored by: BBOX.io Differential Revision: https://reviews.freebsd.org/D59330
A v6 raw socket may ask the kernel to validate the checksum of an inbound packet. If it does, and the validation fails, we discard the packet, but this isn't really right: other raw sockets may wish to receive a copy of the packet anyway. Rework checksum handling to address this problem, and use a flag to avoid computing the checksum more than once for a given packet. Fixes: https://cgit.freebsd.org/src/commit/?id=de2d47842e880281 ("SMR protection for inpcbs") Reviewed by: pouria, glebius Reported by: Yunzhi Ke MFC after: 1 week Sponsored by: The FreeBSD Foundation Differential Revision: https://reviews.freebsd.org/D58559
pfsync interfaces do not have ifp->if_inet6 set, so when we update the
MTU for those interfaces we panicked.
Add an explicit check for this. This should be temporary, until pfsync
is no longer a struct ifnet (as we've already done for pflog).
Reviewed by: glebius
Sponsored by: Rubicon Communications, LLC ("Netgate")
Differential Revision: https://reviews.freebsd.org/D58701
pfsync packets were allocated with m_get2(), which can't return packets larger than MJUMPAGESIZE. As a result 9k MTU pfsync interfaces simply didn't work. Use m_get3(), which can allocate sufficiently large mbufs. Extend the pfsync:bulk test case to provoke this problem. PR: https://bugs.freebsd.org/bugzilla/show_bug.cgi?id=297307 MFC after: 2 weeks Sponsored by: Rubicon Communications, LLC ("Netgate")
ifconfig: Add SR-IOV VF status output - Adds SR-IOV VF status to the existing ifconfig "-v" output - Adds ioctl command for reporting VF status info from drivers - Adds support to iflib for drivers to handle this new ioctl - Add support for ioctl in ixl(4) Signed-off-by: Eric Joyner <erj@freebsd.org> Relnotes: yes Differential Revision: https://reviews.freebsd.org/D19647
iflib: Avoid locking for unsupported VF status queries ifconfig -v requests SR-IOV VF status from every interface. iflib previously acquired the context lock before dispatching the request even for VFs and drivers using the default unsupported method. Mailbox work on a VF could therefore delay the complete interface listing. VF status describes the children of an SR-IOV PF. Reject requests on VF contexts and classes using the default method without taking the context lock. Keep the lock for actual PF status providers. Fixes: https://cgit.freebsd.org/src/commit/?id=1ccf543b21ef ("ifconfig: Add SR-IOV VF status output")
Replace the records with a versioned nvlist transported through struct ifreq, following SIOCGIFCAPNV. The network stack now packs and copies results, supports bounded retry for larger results, and handles native and 32-bit callers centrally. Drivers only populate a kernel nvlist while their state is locked. Define optional common fields for identity, configuration and handshake state, VLAN policy, queue resources, runtime blocks, PF link state, and namespaced driver extensions. Document the extension and versioning contract and require providers to omit values they cannot observe. Improve the ixl provider to track its mailbox handshake and report the expanded common policy. Render the expanded status as grouped output under ifconfig -v.
Honor RTEXT_FILTER_VF on RTM_GETLINK requests and expose the versioned SR-IOV VF status through typed nested FreeBSD attributes. Report IFLA_NUM_VF with a successful requested query and preserve per-provider errors in the status container. Map the common nvlist schema to native integer, boolean, string, and binary attributes. Carry namespaced driver extensions as packed versioned nvlists so adding a driver-specific field does not expand the common netlink ABI. Add SNL parsers, parser verification, a constructed nested-status test, and an RTM_GETLINK test for an interface without SR-IOV support. Document the query contract and every attribute.
The ifdi_init method cannot report an error, so iflib always marks an interface running and enables its interrupts after the callback returns. Drivers whose hardware initialization depends on an unavailable peer can only return early and leave a falsely running interface. Add iflib_init_failed() so a callback can leave the interface stopped. Also add a conditional reset request for asynchronous recovery: it is discarded if the interface is administratively down when the admin task runs, preventing a queued retry from resurrecting a stopped interface. Do not restore saved driver flags after an MTU or capability change when initialization failed. Restoring the pre-init flags would overwrite the stopped result with stale RUNNING state. Document that reset requests require the caller to schedule the admin task, that output remains blocked during recovery, and that iflib rather than the driver owns the driver flags. MFC after: 2 weeks
When crypto_dispatch() or crypto_dispatch_async() returns non-zero, the registered callback is never invoked. In both ovpn_transmit_to_peer() and ovpn_udp_input(), if_ovpn.c did not free the cryptop request, release the peer/sc reference count, or free the mbuf on dispatch failure. This results in three simultaneous leaks per failed dispatch: - crp allocated via crypto_getreq() is never freed - peer->refcount (encrypt) or sc->refcount (decrypt) incremented but never decremented - mbuf passed to crypto_use_mbuf() is never freed The leaks are reachable under memory pressure when the OCF scheduler returns ENOMEM from crypto_dispatch(). The registered callbacks (ovpn_encrypt_tx_cb, ovpn_decrypt_rx_cb) correctly handle crp_etype for crypto operation failures; this fix addresses the separate dispatch-level failure path where no callback is invoked. Found during code review following FreeBSD-SA-26:52.if_wg. Reviewed by: kp Differential Revision: https://reviews.freebsd.org/D58754
10GBase-BX uses paired wavelengths to carry both directions over a single strand of single-mode fiber. The optics must be paired so that the transmit and receive wavelengths cross over. MFC after: 2 weeks
Add missing TF_DISCONNECTED bit. Reported by: Hannes Elfert Fixes: https://cgit.freebsd.org/src/commit/?id=40dbb06fa73c ("inpcb: retire INP_DROPPED and in_pcbdrop()")
Initialize all three hashes (exact, wild, load balance group) with a per- bucket lock. Nothing changes for the packet lookup KPI - it still uses SMR section for thread safety. But connect(2) and bind(2) operations gain parallelism now. The main concept is that as we lookup inpcb database for editing, we are accumulating bucket locks necessary to accomplish the operation. Once all lookups are complete and we are good to go, the inpcb is inserted (or moved) and accumulated lock context is released. Reviewed by: markj Differential Revision: https://reviews.freebsd.org/D58131
Fixes and new features include: hostapd: * support RSN overriding (e.g., WPA3-Personal Compatibility Mode) * EHT/IEEE 802.11be/Wi-Fi 7 - more complete support - fix message validation issues that could enable DoS attacks - fix group key rekeying * enable SAE group 20 by default if SAE-EXT-KEY is enabled * reject unexpected SAE password identifier to avoid DoS attack against a specific STA * mandate use of SAE H2E when using password identifiers * assign VLAN when using SAE with PMKSA caching * support SPP A-MSDU negotiation * support IEEE 802.11bi functionality - changing SAE password identifiers - EPPKE - IEEE 802.1X/EAP in Authentication frames - Association frame encryption - PMKID privacy * remove the driver interface for now obsolete Host AP driver * remove the driver interface for now obsolete Atheros WEXT interface * move supported, basic, and Beacon TX rate configuration to be at BSS level instead of per-radio for all BSSs * fix various issues in Multiple-BSSID functionality * support OpenSSL 3.0 API changes * EAP-TEAP: protocol changes based on RFC 9930; this is not compatible with previous versions * support Automated Frequency Coordination (AFC) on the 6 GHz band * improve GAS/ANQP processing to support larger ANQP responses * a large number of other fixes, cleanup, and extensions wpa_supplicant: * support RSN overriding (e.g., WPA3-Personal Compatibility Mode) * improve BSS transition management support * EHT/IEEE 802.11be/Wi-Fi 7 - more complete support - fix message validation issues that could enable DoS attacks * support Wi-Fi Direct R2 * support Wi-Fi Aware (add synchronized NAN; extend USD support) * support Proximity Ranging * support SPP A-MSDU negotiation * support IEEE 802.11bi functionality - changing SAE password identifiers - EPPKE - IEEE 802.1X/EAP in Authentication frames - Association frame encryption - PMKID privacy * enable layer 2/Wi-Fi multicast filtering for all networks (not just some Passpoint networks which enabled this before) * wpa_gui: port to Qt6 * support OpenSSL 3.0 API changes * EAP-TEAP: protocol changes based on RFC 9930; this is not compatible with previous versions * maintain configuration file permissions when writing updated configuration * add option to validate PKCS#11/OpenSC engine and module paths * fix PMKSA caching to enforce network context to avoid misuse of unexpected PMKSA cache entries * fix a potential DoS attack in SAE processing of an unexpected element * fix incomplete bounds checking of mesh AMPE messages that could have resulted in DoS attacks and memory corruption * a large number of other fixes, cleanup, and extensions MFC after: 2 months Merge commit '513264698a892550ba20aeb9da7f360417fdc63b'
When the first loop in inm_merge() hits an error, generally because it
hit some limit on the number of source filters for a multicast group,
inm_merge() tries to atomically roll back changes to the group source
filter list.
To roll back, it iterates over the global source filter list for the
multicast group, starting at the last entry that we updated ("nims").
But, if we have not yet updated any entries, this variable is
uninitialized. Initialize it to NULL, so that RB_FOREACH_REVERSE_FROM
doesn't visit any source filters in this case.
All of the above applies to the v6 case.
Reported by: Daniel Birtwhistle
MFC after: 1 week
Sponsored by: The FreeBSD Foundation
led(4) invokes driver callbacks while holding its mutex, including from a callout. iflib_led_func() cannot acquire the sleepable context lock in those contexts without causing a lock-order reversal or sleeping from the callout. Record the latest requested state under the iflib state lock and enqueue the existing per-device taskqueue. The task can safely take the context lock before invoking the driver. Coalescing requests also avoids accumulating stale blink transitions when hardware access is slow. Destroy the LED device before draining its task so no new callback can race driver detach. MFC after: 2 weeks
When a driver implements ifdi_led_func, have the framework create its led(4) device after attach completes and the ifnet and context locks are released. PR: https://bugs.freebsd.org/bugzilla/show_bug.cgi?id=246885 Reported by: jlduran Reviewed by: markj MFC after: 2 weeks Differential Revision: https://reviews.freebsd.org/D32389
A driver class may implement LED control even though the capability is not available on every device or firmware version it supports. Add an optional capability method and consult it before creating the led(4) device. Default to supported so existing providers are unchanged. This will be used by bnxt which blends PF and VF in the same driver. MFC after: 2 weeks
Tables that have one element per protocol or address family were previously sized by AF_MAX + 1 since AF_MAX was off by one. Now that AF_MAX has been corrected, we need to apply the opposite correction to these tables. Fixes: https://cgit.freebsd.org/src/commit/?id=ddd850aa7720 ("sys/socket.h: Fix AF_MAX") MFC after: 3 days Sponsored by: Klara, Inc. Sponsored by: NetApp, Inc. Reviewed by: pouria, kevans, glebius Differential Revision: https://reviews.freebsd.org/D58826
Reviewed by: markj Sponsored by: The FreeBSD Foundation MFC after: 1 week Differential revision: https://reviews.freebsd.org/D58646
for SADB_UPDATE op SADB_X_EXT_NEW_ADDRESS_SRC/DST extensions, by checking the sa_len matching the address family requirements before doing the copy. Also convert KEY_SETSECASIDX() and KEY_SETSECSPIDX() to functions and apply the sa_len clamping there. See https://github.com/0xdeadbeefnetwork/pfkey-sadb-overflow PR: https://bugs.freebsd.org/bugzilla/show_bug.cgi?id=297264 Tested by: Wafa Hamzah <wafah@nvidia.com> (previous version) Reviewed by: markj Sponsored by: The FreeBSD Foundation MFC after: 1 week Differential revision: https://reviews.freebsd.org/D58646
Noted and reviewed by: markj Sponsored by: The FreeBSD Foundation MFC after: 1 week Differential revision: https://reviews.freebsd.org/D58646
- Fixed memory leaks around m_dup() not freeing the original chain on
failure. If we return ENOMEM, we are expected to have freed the
chain, else the mbuf would be leaked. Also updated iflib_ether_pad()
to follow the same structure.
- In iflib_parse_header()
o Fixed a bug where the ip/ip6 and th pointers may point into a
freed chain after m_pullup. Those pointers must be reset to
point into the new chain.
o Eliminate ENXIO returns for non-TCP TSO sends (which would violate
the mbuf ownership contract if they could happen). Since they
cannot happen, I made them assertions instead.
- in iflib_ether_pad(), return ENOMEM after freeing mbuf, so that
mp_ring knows it is free. An ENOBUFS error will cause the mp_ring
path to retain the mbuf and retry
- in iflib_encap():
o Fix a leak when bus_dmamap_load_mbuf_sg() returns ENOMEM
o Fix a use-after-free in the mp_ring path when a driver using
ktls frees an mbuf and returns ENOBUFS via iflib_encap()
After this change the expection from iflib_encap is that:
mp_ring: ENOBUFS can be returned only when we run out of descriptors
(ENOBUFS causes mp_ring to retain the mbuf).
simple_tx: iflib_encap() always consumes the mbuf, regardless of the
return
Note that iflib_debugnet_transmit(), like simple_tx, expects that
iflib_encap() always consumes mbufs. This will be true after mp_ring
is removed, and its such a rare special case (overrunning the ring
during panic dumps) that I don't think its worth fixing in the
meantime.
Sponsored by: Netflix
Reviewed by: kbowling, sumit.saxena_broadcom.com
Differential Revision: https://reviews.freebsd.org/D58843
Fixes: https://cgit.freebsd.org/src/commit/?id=074ff8746388
lookup_route is only called for outgoing traffic, therefore check nh_ifp index instead of nh_aifp as specified by RFC3542 sec 6. Differential Revision: https://reviews.freebsd.org/D58544
Add six device-scoped fail(9) points at the registration milestones needed to exercise each unwind path. An exact, runtime-only device selector prevents unrelated iflib devices from consuming an armed point. Mark the points non-sleepable because registration holds the ifnet and context locks. Document one-shot operation and bus-address reprobe so a failed attach can be recovered without another kernel build. Reviewed by: gallatin MFC after: 2 weeks Sponsored by: BBOX.io Differential Revision: https://reviews.freebsd.org/D58722
iflib_device_deregister() sets IFC_IN_DETACH before removing the interface, but a task which already passed its detach check can still report a link change. This can re-arm if_linktask after ether_ifdetach() has drained it and leave work pending across queue teardown. Drain the entire private taskqueue before ether_ifdetach(). Drivers may register their own link-related configuration tasks there, so draining only the framework admin task leaves the same race for those drivers. MFC after: 2 weeks Differential Revision: https://reviews.freebsd.org/D58452 Co-authored-by: Andrew Gallatin <gallatin@FreeBSD.org> Co-authored-by: Kevin Bowling <kbowling@FreeBSD.org>
Add an exact-device fail point immediately after the admin task checks IFC_IN_DETACH. This makes the detach race reproducible without affecting another interface. Use a bounded delay to keep the task active while detach enters the taskqueue drain. Mark the point nonsleepable as a safety backstop, and document a one-shot test for verifying that deregistration drains an already-running task before ether_ifdetach(). Reviewed by: gallatin, kgalazka MFC after: 2 weeks Sponsored by: BBOX.io Differential Revision: https://reviews.freebsd.org/D58720
The VFLR task was initialized only from drivers MSI-X interrupt assignment paths. ixl's legacy interrupt handler can nevertheless defer VFLR work, leaving an uninitialized task. Even with MSI-X, the admin interrupt was established before the task was initialized. Initialize it alongside the other private tasks. The existing detach check and private-taskqueue drains then cover its lifecycle for every interrupt mode and registration failure. MFC after: 2 weeks Sponsored by: BBOX.io
When getting some baseline ALTQ numbers, I noticed that if simlple_tx is enabled in kenv, we wind up re-setting the transmit routine, but I neglected to actually clear ctx->ifc_sysctl_simple_tx. That leads to many different panics as we run a mixture of mp_ring and simple_tx. Pointy-hat to: gallatin Sponsored by: Netflix
Allow disabling RXCSUM and RXCSUM6 on an epair interface. If disabled, epair unsets the mbuf flags that indicate a valid checksum when transferring a packet from one epair end to the other. This gives a user in a jail the power to control whether the user wants to use the result of a previous validation (by a physical interface) or not. Reviewed by: kp, tuexen MFC after: 1 month MFC to: stable/15 Differential Revision: https://reviews.freebsd.org/D58786
When taking a snapshot of the before ip_len (s1) for comparison with the after-translated ip_len (s2), we must convert it from network to host byte order before we can use it. Add the missing ntohs() call. PR: https://bugs.freebsd.org/bugzilla/show_bug.cgi?id=296944 MFC after: 3 days
net80211: fix WEP transmit This was broken in 2022 with a security fix (61605e0ae5d8f) which disallowed defaulting to the default TX key if there's no unicast key. Unfortunately this path was also used by WEP transmit. To fix it, add a separate check which ensures that WEP is configured (authtype OPEN, privacy enabled) - then also check if the default TX key is set and that said key is a WEP key. Fixes: https://cgit.freebsd.org/src/commit/?id=61605e0ae5d8f Locally tested: * rtwn(4) AP and rtwn(4) STA w/ static WEP keys configured Differential Revision: https://reviews.freebsd.org/D58854
net80211: add key get/set methods Introduce net80211 key get and set methods with appropriate bounds checking and buffer zero'ing. Differential Revision: https://reviews.freebsd.org/D58705
net80211: migrate the ioctl API to a 128 bit specific API + use key API * Begin migrating the ioctl code to use the key management APIs. Not all of it has been migrated (notably the WEP API hasn't.) * Take special care to copy the TKIP MIC in and out correctly. * Note that some of the defines used as sizes are actually the ioctl sizes, they'll need to be fixed before I push this into a review. * Document this current API as a specific 128 bit key + 128 bit TKIP MIC API. The goal here is to solidify this stuff as the 128 bit ioctl API and not change it, even if net80211 will eventually grow 256 and 384 bit key support. Notably the TKIP stuff - the driver_bsd.c code puts the TKIP after the normal key contents, whereas the net80211 code puts the TKIP stuff in the /end/ of the key buffer. They happen to be equivalent when ioctl key buffer size == net80211 key buffer size, but as I learnt the last couple times I tried this, they're not always going to be equivalent. Differential Revision: https://reviews.freebsd.org/D58384
Add another HE define needed by the upcoming espwl(4). MFC after: 3 days
Modern Netlink arrays encode their elements as repeated attributes of the same type rather than as children of an additional array container. Add an SNL callback that parses one nested element for each occurrence and appends it to a geometrically grown parser array. Retain the existing parray callback for protocols that use the legacy container form. Store the growth capacity in struct snl_parray, appended after its existing public count and items fields so their offsets remain stable on LP64 and ILP32. Require parser targets to be real snl_parray objects, and convert bitset, generic Netlink, and route multipath arrays accordingly. This avoids relying on layout aliases for private growth state. Add regression coverage for a nested bit array that grows beyond its initial allocation, while preserving replacement semantics when a legacy array target is reused. Reviewed by: melifaro, pouria Sponsored by: BBOX.io Differential Revision: https://reviews.freebsd.org/D58775
Add support nexthop statistics and update its manual. While here, fix manual of other nexthop related options. Reviewed by: kfv Differential Revision: https://reviews.freebsd.org/D58538
radix_lockless algorithm creates its own radix tree and allocates its own radix_masks by directly calling rnh_addaddr(). However, during destruction, it only frees the radix_tree without freeing its allocated radix_masks. Fix the leak by calling rn_delete() during radix_destroy(). PR: https://bugs.freebsd.org/bugzilla/show_bug.cgi?id=297339 Reviewed by: melifaro MFC after: 2 weeks Differential Revision: https://reviews.freebsd.org/D59112
Include the entries of the BBLOG STATE column in the computation of the column width as done also for other columns. Also set first to falso only when needed. MFC after: 3 days MFC to: stable/15 MFC to: stable/14 Sponsored by: Netflix, Inc.
pfr_create_kentry() can return NULL. Don't dereference the pointer it returns until after we've checked it. Fixes: https://cgit.freebsd.org/src/commit/?id=08ed87a4a276 ("pf: convert DIOCRSETADDRS to netlink") See also: https://redmine.netgate.com/issues/23622 Sponsored by: Rubicon Communications, LLC ("Netgate")
Provide a convenient in-place sed edit command to update the FreeBSD VersionAddendum dates with today's date. Sponsored by: The FreeBSD Foundation
Several cleanups in tcp_input_with_port(): * Don't assign m twice. * Don't reassign pointers without having done pullup(). * While there, change the type of isipv6 to bool, since it is used that way. No functional change intended. Reported by: Hannes Elfert Reviewed by: pouria, Timo Völker, Nick Banks MFC after: 1 week MFC to: stable/15 Differential Revision: https://reviews.freebsd.org/D59142
pf states may be looked up using one of two keys: the stack key or the
wire key. For states involving address translation, these will be
distinct; the stack key describes the addresses seen by the local
network stack, and the wire key has the translated addresses.
Historically, pf would avoid allocating separate keys if both are
identical. This changed in commit fcdb520c1b4e ("pf: nat64") to always
allocate separate state key structures. Incidentally, OpenBSD seems to
maintain the optimization, but also has an explicit reference count
embedded in state keys.
The change breaks another optimization: pf_state_key_attach() still uses
state key pointer equality to check whether the stack and wire keys are
equal, so those checks are always false after the aforementioned commit.
Thus we never skip the second key lookup, even when that's possible
(i.e., no address translation is involved).
So, for some rulesets we're consuming more memory than needed and
performing more state key lookups than needed. The behaviour of always
looking up the stack key also happens to break some existing rulesets
involving RDR and divert-to, which is how I noticed the problem. I
think those rulesets effectively worked by accident before, but it seems
worth restoring the optimization regardless.
Reviewed by: kp
MFC after: 2 weeks
Fixes: https://cgit.freebsd.org/src/commit/?id=fcdb520c1b4e ("pf: nat64")
Sponsored by: OPNsense
Sponsored by: Klara, Inc.
Differential Revision: https://reviews.freebsd.org/D58922
Use ifa_ifwithaddr_fib() to fix fib-specific ifa lookups with route. Differential Revision: https://reviews.freebsd.org/D59128
When using -b, provide in addition to the BBLog state also the number of entries stored at the endpoint, the corresponding upper limit and the next sequence number to be assigned. Reviewed by: rrs MFC after: 1 week MFC to: stable/15 MFC to: stable/14 Sponsored by: Netflix, Inc. Differential Revision: https://reviews.freebsd.org/D59131
Netlink IFLA_GROUP works with a single group id, in our implementation an interface can be joined to multiple groups and it works with group name. Store interface groups in IFLAF_GROUP attribute. Reviewed by: glebius, melifaro Discussed with: markj Differential Revision: https://reviews.freebsd.org/D58643
Commit 8572367b6814 ("pf: remove STATE_LOOKUP") introduced two seemingly
unintentional changes with respect to divert(4)-injected packets (i.e.,
the PACKET_LOOPED case): we no longer return the matching state, and
direct callers of pf_find_state() now treat matches of diverted packets
the same as having no matching state at all.
This seems inadvertent, and breaks certain rulesets which use divert-to.
Fix them, and add a regression test case.
Fixes: https://cgit.freebsd.org/src/commit/?id=8572367b6814 ("pf: remove STATE_LOOKUP")
Reviewed by: kp
MFC after: 2 weeks
Sponsored by: OPNsense
Sponsored by: Klara, Inc.
Differential Revision: https://reviews.freebsd.org/D59015
Otherwise close() fails. PR: https://bugs.freebsd.org/bugzilla/show_bug.cgi?id=297977 Reviewed by: glebius Fixes: https://cgit.freebsd.org/src/commit/?id=ea7be1293b48 ("keysock: do not use raw socket code") MFC after: 1 week
pf sends outbound packets by offloading them to a single per-vnet SWI handler through the `V_pf_sendqueue` mbuf queue. A large DDoS attack may overwhelm that per-vnet queue with syncookie packets and cause contention in the SWI handler that negatively affects other pf operations. Fix this by sending the initial syncookie challenge from the context of the receiving thread. This avoids the syncookie-induced contention on the `pf_intr` mbuf queue. Sponsored by: Klara, Inc. Sponsored by: Entersekt MFC after: 3 weeks Reviewed by: kp Differential Revision: https://reviews.freebsd.org/D59068
getifaddrs(3) may return AF_INET6 addresses even when netstat(1) is built without INET6 support. Avoid passing these addresses to process_ifa_addr(), which does not handle AF_INET6 in that case. This fixes corrupted column width calculations and runaway output from netstat -i, which could cause periodic daily check output to generate multi-gigabyte mail messages and exhaust disk space and memory. MFC after: 2 weeks
In bridge_input, sc is initialized to NULL and doesn't get resolved until after the Ethernet header pullup. So the pullup's failure path ends up dereferencing the NULL sc when bumping up IFCOUNTER_IERRORS. The m_freem call right under it is redundant as the failure path in m_pullup already freed the chain. Drop both lines, matching what we have in bridge_output. ether_input_internal() discards frames shorter than ETHER_HDR_LEN before the bridge hook, so it is unlikely that it will fire. We still keep the guard as lagg(4) and ng_ether(4) may replace the mbuf before the bridge hook. Signed-off-by: Aaron Espinoza <acesp25@freebsd.org> Reviewed by: pouria Pull Request: https://github.com/freebsd/freebsd-src/pull/2393
For non listening TCP endpoints, increment the tcps_sig_err_sigopt counter when TCP MD5 is not enabled in the TCP connection, but a segment containing a TCP MD5 option is received. Also increment the counter when using the RACK or BBR stack. Reported by: Hannes Elfert Reviewed by: rrs MFC after: 1 week MFC to: stable/14 MFC to: stable/15 Differential Revision: https://reviews.freebsd.org/D59249
When _task_fn_admin() is active, it will regularly call IFDI_UPDATE_ADMIN_STATUS(). So there is no need to do it in iflib_media_status. This can be fairly expensive on some drivers (long DELAY busywait loops waiting for a NIC command), and there is no need to pause a userspace app in this DELAY() if it is happening asynchronously anyway. Note the logic to detect if _task_fn_admin() is regularly calling IFDI_UPDATE_ADMIN_STATUS() was copied from that function. Reviewed by: erj, kbowling Sponsored by: Netflix Differential Revision: https://reviews.freebsd.org/D54096
ng_bridge(4) says the node does not learn MAC addresses on uplink hooks. However, learnMac was only checked when inserting a new host. A host already known on a link hook was still moved if a packet with that source address arrived on an uplink hook. The nature of this is that inbound unicast to that host then never arrives (the destination is known on the incoming hook). Unknown unicast after timeout is still sent only to uplink, so the host is not re-learned. The interface stays up and outbound may still work. This can last minutes or weeks until reboot or NGM_BRIDGE_MOVE_HOST. Connecting ng_ether(4) lower to an uplink hook is enough: the host's own transmit can appear on the uplink and the table entry moves. Use the same learnMac test for data-path move as for insert. NGM_BRIDGE_MOVE_HOST from userland is unchanged. MFC after: 1 week Reviewed by: jlduran Differential Revision: https://reviews.freebsd.org/D58902
The restored watchdog arms when the outstanding descriptor count grows, but then continues counting based only on the queue remaining frozen. A single growth sample can therefore leave a quiet, nearly empty queue armed until the watchdog resets the interface. Lockless sampling of the queue counters can also manufacture the initial growth sample. This matches watchdog reports from I354 queues with 979 or 980 of 1022 usable descriptors still available. Neither queue was under transmit backpressure when the reset flapped its link. Keep the watchdog armed only while the outstanding count continues to grow, the software ring is stalled, or the hardware ring is at iflib's backpressure threshold. The last condition preserves hang detection with simple-TX, which bypasses the software ring. A busy hang still reaches the verdict while a frozen but quiet tail disarms. Retain the final driver completion peek so a missed completion interrupt schedules the queue task instead of resetting it. Validated on an 82580 with one and four queue sets in the default mp_ring and simple-TX modes. Sustained traffic and repeated burst/idle cycles produced no false resets. Sixteen-flow runs exercised all four queues in both modes. Clearing TCTL.EN under load in each configuration filled the rings; the reset counter advanced once per injection, reset restored TCTL and the link, and traffic recovered. Tested by: glebius Reviewed by: iflib (gallatin), manpages (ziaee) Fixes: https://cgit.freebsd.org/src/commit/?id=69c3e0de01c1 ("iflib: restore TX watchdog functionality") MFC after: 6 days (after 69c3e0de01c1) Sponsored by: BBOX.io
iflib_device_register() acquired IFNET_WLOCK to preserve lock order when ether_ifattach() was called with the context lock held. The context lock is now released around ether_ifattach(), making registration-wide ifnet serialization unnecessary. Keeping IFNET_WLOCK across driver attachment also allows synchronous interface event handlers to recurse on it. The rtnetlink interface-group dump does so through if_foreach_group() while handling the interface attachment event. Remove the outer lock and the corresponding failure-path unlock and relock transitions. Continue to drop the context lock around ether_ifattach() and taskqueue drains, and preserve context-lock coverage for driver attach and detach. Validated under WITNESS on 82576 and I226 controllers. Multiple VF attach and detach cycles, netmap control operations, and every iflib registration failure injection point completed without lock or cleanup errors. PR: https://bugs.freebsd.org/bugzilla/show_bug.cgi?id=298121 Reported by: glebius, netchild, Yuichiro NAITO <naito.yuichiro@gmail.com> Reviewed by: gallatin, glebius Fixes: https://cgit.freebsd.org/src/commit/?id=e0e12405285b ("netmap: fix LOR in iflib_netmap_register") Fixes: https://cgit.freebsd.org/src/commit/?id=2f8f892ca344 ("rtnetlink: Add FreeBSD-specific IFLAF_GROUP support") Fixes: https://cgit.freebsd.org/src/commit/?id=90e7dbe5e2ca ("iflib: Add registration failure injection points") MFC after: 2 weeks Sponsored by: BBOX.io Differential Revision: https://reviews.freebsd.org/D59294
The new watchdog code triggers spurious watchdog resets on NICs doing KTLS offload. Fix this by using the actual segments consumed by the NIC driver's isc_txd_encap. The issue is that rs_pending is updated using an estimate of the descriptors that will be used for the current packet, based on what bus_dma produced. However, NICs which support ktls offload may do extra DMAs (and consume extra descriptors) to derive crypto state when re-transmitting TLS segments. This is the reason for allowing drivers to control ift_pad. When this happens, the estimated rs_pending may undercount. This may also happen if NIC drivers consume extra descriptors for other reasons. (eg, hw errata handling on e1000) Reviewed by: kbowling Differential Revision: https://reviews.freebsd.org/D59321 Sponsored-by: Netflix
When receiving a SYN segment with an MD5 option on a listening socket, which has not enabled TCP MD5 support, increment the counter for unexpected signatures (tcps_sig_err_sigopt). Reviewed by: rscheff MFC after: 1 week MFC to: stable/14 MFC to: stable/15 Differential Revision: https://reviews.freebsd.org/D59303
When processing the ACK of the initial TCP handshake using the SYN cookie, don't increment the counter for unexpected signatures (tcps_sig_err_sigopt). The correct counter (tcps_sig_rcvbadsig) for bad signatures is already incremented in TCPMD5_INPUT(). Reported by: Hannes Elfert Reviewed by: rscheff MFC after: 1 week MFC to: stable/14 MFC to: stable/15 Differential Revision: https://reviews.freebsd.org/D59302
All other usages of SCF_SIGNATURE are protected by IPSEC_SUPPORT or TCP_SIGNATURE. No functional change intended. Reported by: Hannes Elfert MFC after: 1 week MFC to: stable/14 MFC to: stable/15
Approved by: kp Sponsored by: InnoGames GmbH Differential Revision: https://reviews.freebsd.org/D58755
Add a transport neutral kernel snapshot for NIC-specific SR-IOV VF status and an optional iflib provider method. Providers gather state under driver defined synchronization. Honor RTEXT_FILTER_VF on RTM_GETLINK requests and encode the status as native typed route Netlink attributes. Represent VFs, driver namespaces, and namespace fields as directly repeated nested attributes. Presence masks in consumers can distinguish omission from false or zero. Drivers may add custom status under stable, versioned namespaces. The named, typed representation lets generic transports and consumers carry or display fields without knowing their driver-specific schemas, while the driver retains ownership of their names and meanings. Document the ABI and add parser and RTM_GETLINK coverage. Reviewed by: melifaro, iflib (gallatin), kgalazka (previous version) Sponsored by: BBOX.io Differential Revision: https://reviews.freebsd.org/D58776
Since the struct inpcb in embedded in the struct tcpcb, the relationship can't change. So there is no need to reassign the tp anymore. No functional change intended. Reported by: Hannes Elfert Reviewed by: glebius MFC after: 1 week MFC to: stable/15 Differential Revision: https://reviews.freebsd.org/D59384
As per RFC 3590 MLD Report and Done messages are permitted to use the
unspecified address as a source address (e.g. during duplicate address
detection for the first IPv6 address). Allow this, but only this.
Reported by: Alexander Leidinger <Alexander@Leidinger.net>
Reviewed by: bms
See also: OpenBSD, sashan <sashan@openbsd.org>, 60036e8507
Sponsored by: Rubicon Communications, LLC ("Netgate")
Differential Revision: https://reviews.freebsd.org/D59334
Revision 1.1212 of pf.c weakened the TCP reset check in stateful
connection tracking to let legitimate resets pass in the backwards
window. Such a reset is accepted only if its acknowledgment number
matches perfectly. But as a workaround for broken stacks, pf
replaces an acknowledgment number of 0 in a reset with the tracked
sequence of the peer. Then the perfect match always succeeds, and
an attacker can spoof resets more easily than intended. Use the
acknowledgment number from the wire, before the workaround has
modified it.
discovered by Minghao Zhang; OK sashan@
Obtained from: OpenBSD, bluhm <bluhm@openbsd.org>, 1e0a1f4b82
Sponsored by: Rubicon Communications, LLC ("Netgate")
when collecting necessary locks for the operation. The IPv6 version is already correct. Submitted by: Nick Price <nick spun.io> Fixes: https://cgit.freebsd.org/src/commit/?id=1dda8ba77a20d983774a3bf0e94c51c93f1d8758
The syncache entry holds one TCPS_SYN_RECEIVED count that normally is transferred to the the newborn tp. Upon failure syncache_socket() shall not use TCPSTATES_INC/TCPSTATES_DEC (see 5050df3f4aa4 why). But when syncache_socket() fails in_pcbconnect(), it calls tcp_discardcb() to free resources that were just allocated by tcp_newtcpcb() and this tcp_discardcb() would do TCPSTATES_DEC(tp->t_state). The t_state is TCPS_CLOSED at this point. Make tcp_discardcb() symmetrical to tcp_newtcpcb() - not responsible for the TCPSTATES. Make the caller responsible for state count book keeping. Reviewed by: tuexen Fixes: https://cgit.freebsd.org/src/commit/?id=3703e1a73e0e0367c04f47f793e46495e46e647b Differential Revision: https://reviews.freebsd.org/D59325
PF-administered priority need not impose an access VLAN. Permit the existing VLAN PCP attribute in either access or trunk mode, while keeping VLAN identifier and protocol access only. Display trunk PCP in ifconfig and document the broader meaning in the kernel snapshot, libifconfig, and rtnetlink contracts. This changes no attribute numbers, wire encoding, or structure layout. Reported by: kib Sponsored by: BBOX.io
ND_OPT_ROUTE_INFO is 24, so an access of that index is technically out of bounds, though note that the per-option struct has an entry for this option. Fix the array size to appease -fsanitize=array-bounds. Reviewed by: pouria, glebius MFC after: 2 weeks Sponsored by: The FreeBSD Foundation Differential Revision: https://reviews.freebsd.org/D59392
This fixes the build with gcc 16 after the changes to -Wunused*[0]. [0] https://gcc.gnu.org/gcc-16/porting_to.html#changes-to-wunused Reviewed by: cy MFC after: 3 days Sponsored by: The FreeBSD Foundation
This fixes the build with gcc 16 after the changes to -Wunused*[0]. [0] https://gcc.gnu.org/gcc-16/porting_to.html#changes-to-wunused Reviewed by: gallatin MFC after: 3 days Sponsored by: The FreeBSD Foundation Differential Revision: https://reviews.freebsd.org/D59544
sbuf_new_auto() sleeps, and ng_hci_default_rcvmsg() can run under the raw HCI socket pcb mutex held across NG_SEND_MSG_PATH(). Reported by: WITNESS Fixes: https://cgit.freebsd.org/src/commit/?id=ad91d47db306 Reviewed by: glebius, adrian Differential Revision: https://reviews.freebsd.org/D59550
This fixes the build with gcc 16 after the changes to -Wunused*[0]. [0] https://gcc.gnu.org/gcc-16/porting_to.html#changes-to-wunused Reviewed by: tuexen MFC after: 3 days Sponsored by: The FreeBSD Foundation Differential Revision: https://reviews.freebsd.org/D59590
bridge_pfil() returned a fragmentation failure without counting it, and bridge_fragment() dropped a chain on three allocation failures without counting those either. Count the first on the filtered interface and the others with ips_odropped, which is what ip_fragment() uses for the same failure and what bridge_fragment() already uses for its success case. Reviewed by: gallatin Differential Revision: https://reviews.freebsd.org/D59391 Assisted-by: Claude Code (Fable 5, Opus 5)
Application may set sin_port to some value. Although this value is not used by SOCK_RAW, it breaks a check that the address is available. A perfect fix would be not use sockaddrs for ifaddrs, but that would be a bigger change. PR: https://bugs.freebsd.org/bugzilla/show_bug.cgi?id=298366 Reviewed by: pouria, bnovkov, adrian Differential Revision: https://reviews.freebsd.org/D59570 Fixes: https://cgit.freebsd.org/src/commit/?id=948ad32ae1e0811f45e1d38f26636fefed5051f0
Implement buf_ring/drbr deferred transmit in iflib. This is intended to allow the new simpler code path to replace mp_ring. This patch makes the simple_tx outperform mp_ring by a wide margin when CPU is the bottleneck (eg, cannot fill the NIC). See graphs at: https://people.freebsd.org/~gallatin/mpring_vs_simple_tx Note that the buf ring is used for contention, not capacity. Eg, it is used as a place for contending threads to put packets without waiting for a mutex. It is not designed to act as a software ring on top of the hardware descriptors provided by the underlying NIC driver. "stranded packets" are exceedingly rare due to the fact that if there is enough load to use the buf_ring, there will probably be more load coming that can be a drainer. Not scheduling a gtask to drain is intentional, and we really on the timer as a fallback. One thing I noticed while developing this patch is that a simple mutex with no deferral generally outperformed both mp_ring and drbr at high levels of contention for the same queue. This is inherent in a bounded MPSC queue where multiple producers are contending on claiming ring entries. So I came up with the idea of bounding the number of producers such that the deferral ring would devolve to a mutex when contention was high. Identifying the crossover point in a general way was hard. On different machines, the point between a mutex and a deferral ring was very different and also depended on the placement of the producers. I eventually realized that on the large AMD EPYC servers that I was testing on, the crossover point generally coincided with the work spilling into another CCX. So I developed an approach where we limit the number of producers by limiting simultanious producers to the same L3 AND by bounding the number of simultanious producers. This limit can be adjusted via the sysctl net.iflib.max_producers. The number of packets drained from the deferral ring is limited by net.iflib.simple_drain_quota. The intent is to process just enough packets in the gtaskq context so as to make space to allow threads to make progress. The task also uses a trylock so as to avoid blocking, waiting for a thread to drain. This change also enables ALTQ support for simple tx. Reviewed by: kbowling Sponsored by: Netflix Differential Revision: https://reviews.freebsd.org/D58901
Use the context-locked datapath state for netmap initialization results, resume completion and administrative work eligibility. Keep stopped and failed interfaces eligible for link and mailbox recovery without using IFF_DRV_OACTIVE as an implicit indication that initialization has been attempted. Evaluate media and deferred admin eligibility under the context lock rather than from an earlier IFF_DRV_ flag snapshot. Always pass through iflib_stop() before an IOV-init callback which changes the PF queue layout, as the IOV-uninit counterpart already does. Neither administrative down nor a cleared IFF_DRV_RUNNING proves that queue DMA is stopped. The private lifecycle state decides whether the hardware stop is necessary; IFF_UP only decides whether to initialize again afterward. Reviewed by: iflib (gallatin) MFC after: 2 weeks Sponsored by: BBOX.io Differential Revision: https://reviews.freebsd.org/D59597
Publish an atomic software run state snapshot through iflib_is_running(). Use it for iflib traffic admission, queue tasks, timers, debugnet and live-configuration checks instead of reading the unsynchronized ifnet driver flags. Keep admission closed after failed initialization and close it when a watchdog requests deferred recovery. Retain the context-locked datapath state for DMA ownership: a closed admission gate does not establish that the hardware is stopped. Open the gate after receive buffer setup and before interrupt enable, at the existing RUNNING publication point. Serialize the writers with the state mutex and continue publishing RUNNING/OACTIVE for network stack consumers. The accessor takes no lock and is usable from filters, but is only a snapshot, not a context reference or a queue-user drain. Recheck multicast and VFLR admission under the context lock. Replace the OACTIVE drain check with the same private admission gate, retaining the separate per-queue descriptor backpressure policy. Rename its debug counter to txq_drain_stopped. Document the accessor and its synchronization limits. This does not remove legacy flag reads elsewhere in the network stack or change the existing queue-drain and device-stop contracts. External drivers will still see the IFF_DRV_ publications but should migrate to this API. Reviewed by: iflib (gallatin) MFC after: 2 weeks Sponsored by: BBOX.io Differential Revision: https://reviews.freebsd.org/D59598
Initialization publishes RUNNING and enables interrupts before clearing the simple-TX producer barrier. A transmitter admitted in that window still sees IFLIB_TXQ_QUIESCING and drops its packet with ENETDOWN. A TX completion task can likewise return without draining deferred packets. Clear the barrier on every TX queue after driver initialization and RX buffer setup succeed, before publishing admission. Keep failed init paths closed and retain the existing stop, interrupt and timer ordering. This changes only initialization; no work is added to the packet path. Reviewed by: iflib (gallatin) Sponsored by: BBOX.io Differential Revision: https://reviews.freebsd.org/D59690
We have a problem in UMA where creating new zones requires a traversal
of all existing zones (in order to find a unique name for the vm.uma
sysctl subtree). This means that per-VNET UMA zones can be expensive to
create if one creates many VNET jails in a row. This itself is a
problem, but I don't see a quick solution.
Ideally we would avoid creating per-VNET zones in the first place
(except when the desire to impose per-VNET limits necessitates this),
and it turns out that the TCP fastopen and SACK code create several of
these for no apparent reason.
So: make fastopen and SACK zones global. Move some fastopen structure
definitions into tcp_fastopen.c, as they don't need to be public.
Reported by: bapt ("creating many VNET jails in a row is slow")
Reviewed by: tuexen, glebius
MFC after: 3 weeks
Differential Revision: https://reviews.freebsd.org/D59708
Until now, simple tx has used its own queue selector, and has ignored isc_txq_select and isc_txq_select_v2 (not this only seems to matter for ice(4) with dcb enabled). This change makes simple-tx use isc_txq_select* when present. The implementation is defined to be efficient, with a transmit routine chosen up-front that hard-codes the queue selection and calls an always-inlined body. This avoids a useless test per packet in the hotpath, and also may avoid speculation into header parsing. Note that this was designed for readability and efficiency in the common case (interface up, not ALTQ). That's why we do queue selection without duplicating nic-running and altq checks, leaving them to the common implmentation. Sponsored by: Netflix Reviewed by: kbowling Differential Revision: https://reviews.freebsd.org/D59712
net.iflib.prefer_mpring can be used to control whether or not all iflib driver instances default to mp_ring or simple_tx. This is intended to be temporary, to allow easy testing (now) of simple_tx, and to allow an easy fallback to the legacy path (later) once the default is switched to simple_tx
Changing the VID of an existing VLAN interface rehashes the interface and announces the new VID, but does not unregister the old VID. Parent drivers and VLAN event consumers can consequently retain stale filter membership. After successfully inserting the new VID, emit vlan_unconfig for the old VID before the existing vlan_config notification. Do not unregister anything if insertion fails and the old VID is restored. MFC after: 2 weeks Sponsored by: BBOX.io
When we receive an ICMP redirect, rib_add_redirect() is used to apply
the redirect to all FIBs. This has been the case since support for
multiple FIBs was added. However, it seems rather dubious: the new
gateway might not be routable from all FIBs, and the validation done for
v4 redirects in icmp_verify_redirect_gateway() is only applied to the
FIB from which the redirect originated.
Modify the handler to apply the redirect only in the originating FIB.
Reported by: Yuxiang Yang, Yizhou Zhao, Ao Wang, Xuewei Feng, Qi Li,
and Ke Xu from Tsinghua University using GLM-5.1 from Z.ai
Reviewed by: pouria, zlei, glebius, melifaro
MFC after: 3 weeks
Sponsored by: The FreeBSD Foundation
Differential Revision: https://reviews.freebsd.org/D59567
fib_algo indexes its idx->nhop array by the nexthop index with assumption of its uniqueness. Which is true except for IPv4 over IPv6 nexthops. Give each index space its own segment within the same array and offset the index by the segment base. Segments are created on demand and sized independently, so the rib's own family keeps base 0 and tables without cross-family nexthops index exactly as before. Reviewed by: melifaro Discussed with: markj MFC after: 2 weeks Differential Revision: https://reviews.freebsd.org/D59552
Not to struct pfioc_table.
MFC after: 1 week
Sponsored by: Rubicon Communications, LLC ("Netgate")
tcp: collect and report more statistics on the TCP host cache Reviewed by: tuexen Differential Revision: https://reviews.freebsd.org/D59451
tcp: turn on TCP hostcache for socket buffer sizes The code has been there since introduction of the TCP hostcache in 97d8d152c28bb in 2003, but was turned off. Reviewed by: tuexen, cc Differential Revision: https://reviews.freebsd.org/D59452
tcp: use sparse initializer for host cache metrics Reviewed by: tuexen Differential Revision: https://reviews.freebsd.org/D59453
Dispatch SIOCGIFRSSKEY and SIOCGIFRSSHASH through typed get_rss_key and get_rss_hash methods under the context lock. Drivers can report their programmed key and hash selections. The default methods return EOPNOTSUPP. The common RSS key does not describe which hash selections a particular device actually programs. hn(4) needs the effective VF configuration when synchronizing RSS between its synthetic and VF receive paths. Reviewed by: iflib (gallatin) MFC after: 2 weeks Sponsored by: BBOX.io Differential Revision: https://reviews.freebsd.org/D59691
Rest of multicast related kernel structures live in in6?_var.h and this piece of ip6?_var.h was anyway requiring in6?_var.h to be included before. Differential Revision: https://reviews.freebsd.org/D55963
Differential Revision: https://reviews.freebsd.org/D59140
ipfw: refactor macros around m_pullup() In the prologue, where we check if the argument is a memory or an mbuf chain, do not set 'struct ip *ip' pointer. However, set the 'struct ether_header *eh' pointer there and set Etherner header length in 'ehlen'. Side effect of this refactor is that now Layer 2 hooks may send mbufs with ETHERTYPE_VLAN frames. However, current network stack doesn't do that. Write a new PULLUP() macro that would take type of the argument to determine how much to pull. Unlike PULLUP_TO() this macro can take typed pointer. This will allow to get rid of 'void *ulp' and bunch of casting macros in the next change. Use local bool variable to see if we need to unlock upon jump to pullup_failed. Embed pointer update into the branch of the macro, where pointers indeed need an update. Use new PULLUP() macro to pullup initial 'struct ip *ip' and 'struct ip6_hdr *ip6'. This removes max_protohdr sized pullup, that previously tried to pull more than an unmapped mbuf could yield, fixing a bug covered by the testcase sys/netpfil/ipfw/unmapped. Reviewed by: gallatin Differential Revision: https://reviews.freebsd.org/D59426
ipfw: use typed pointers to access network protocols headers where possible The 'void *ulp' is still in action, but where possible prefer a typed pointer. Get rid of associated pre-processor macros. Differential Revision: https://reviews.freebsd.org/D59435
ipfw: cleanup !FreeBSD and !_KERNEL code These are leftovers from last import from Luigi Rizzo in 2013, that were never compiled checked since and lots of changes have had happened since and more changes to come. Let's not pretend ipfw is buildable on something else than FreeBSD nor it can be compiled for userland. Reviewed by: lytboris_gmail.com Differential Revision: https://reviews.freebsd.org/D59427
When performing an unconnected sendto() on a v6 UDP socket in a classic
jail, we were not applying the usual policy of replacing the loopback
addr with the jail's primary IP. Compare with, e.g., udp6_connect() or
the IPv4 udp_send(). Fix that.
Reported by: Yuxiang Yang, Yizhou Zhao, Ao Wang, Xuewei Feng, Qi Li,
and Ke Xu from Tsinghua University using GLM-5.1 from Z.ai
Reviewed by: bz, glebius
MFC after: 2 weeks
Sponsored by: The FreeBSD Foundation
Differential Revision: https://reviews.freebsd.org/D59772
When one pfsync host clears states it informs its peers about this.
While processing such messages, in pfsync_in_clr() we failed to take the
interface name into account.
This meant that if one host cleared states on one interface the peers
would clear all states, not just those on the affected interface.
Actually check for the interface in pfsync_in_clr()
Sponsored by: Rubicon Communications, LLC ("Netgate")
fd is allocated with M_ZERO and fd_num_af never decreases. Therefore, no need for zeroing nhaf_count and nhaf_base here. Fixes: https://cgit.freebsd.org/src/commit/?id=633438224304 ("route/fib_algo: Fix nexthop index ...")
Now fd_ref_nhop() returns zero for cross family routes, Do not schedule nhop references and try to rebuild it immediately for connected and static routes. PR: https://bugs.freebsd.org/bugzilla/show_bug.cgi?id=298733 Fixes: https://cgit.freebsd.org/src/commit/?id=633438224304 ("route/fib_algo: Fix nexthop index ...")
Sponsored by: Rubicon Communications, LLC ("Netgate")
Sponsored by: Rubicon Communications, LLC ("Netgate")
ipf_lookup_iterate() validates iter.ili_unit with if (iter.ili_unit < IPL_LOGALL && iter.ili_unit > IPL_LOGMAX) or alternatively, if (iter.ili_unit < -1 && iter.ili_unit > 7) ipf_lookup_add(), ipf_lookup_delete(), ipf_lookup_stats() ipf_lookup_flush() validate with a || Submitted by calif.io for the OpenAI Patch The Planet program Signed-off-by: Andrew Griffiths <andrew@calif.io> Reviewed by: markj MFC after: 1 week Differential Revision: https://reviews.freebsd.org/D59726
When an on-link prefix becomes detached, the kernel keeps generating new RFC 8981 temporary addresses for that prefix. Fix it by ignoring the detached addresses in regen_tmpaddr(). While here, change its return type to bool. PR: https://bugs.freebsd.org/bugzilla/show_bug.cgi?id=298533 Discussed with: markj MFC after: 3 days Differential Revision: https://reviews.freebsd.org/D60051
Since DIOCRSETADDRS was converted to netlink, pf_handle_table_set_addrs() calls pfr_set_addrs() with a NULL size2, as the netlink interface has no buffer to return the deleted addresses in. pfr_set_addrs() only checked size2 for NULL at the end of the function; with PFR_FLAG_FEEDBACK set it dereferenced it unconditionally first. pfctl sets PFR_FLAG_FEEDBACK when run with -v, so "pfctl -v -t foo -T replace ..." panicked the kernel with a NULL pointer dereference. To reproduce: pfctl -e pfctl -t foo -T add 192.0.2.1 pfctl -v -t foo -T replace 192.0.2.2 Check size2 for NULL before dereferencing it, as is already done at the end of the function. The per-address feedback for added and changed addresses is still copied back as before; only the list of deleted addresses, which the netlink caller has no room for, is skipped. While here, compare size2 against NULL explicitly in the second check as well, per style(9). Add a regression test. Approved by: kp (mentor) Fixes: https://cgit.freebsd.org/src/commit/?id=08ed87a4a276 ("pf: convert DIOCRSETADDRS to netlink") MFC after: 1 week Differential Revision: https://reviews.freebsd.org/D60096
pfr_clr_astats() looks up each address it is given and inserts the entry
it finds at the head of a work queue. If the same address is given more
than once, the entry is inserted twice and the second insertion makes it
its own successor. pfr_clstats_kentries() then walks the queue forever,
with the rules lock held for writing, so packet processing and every
other pf operation in that vnet stop as well. To reproduce:
pfctl -e
pfctl -t foo -T add 192.0.2.1
pfctl -t foo -T zero 192.0.2.1 192.0.2.1
Do as pfr_del_addrs() does: clear pfrke_mark on the entries named, then
queue an entry only the first time it is seen. An address given more
than once is cleared, and counted, once. Validate all addresses before
any entry is touched.
Add a regression test.
Reviewed by: kp
Approved by: kp (mentor)
MFC after: 1 week
Sponsored by: Rubicon Communications, LLC ("Netgate")
Differential Revision: https://reviews.freebsd.org/D60103
The SIOCADDMULTI and SIOCDELMULTI handlers add or delete a link-layer
multicast address from an interface's multicast filter list. The
link-layer address is passed using the ifr_addr field of the request
structure.
struct ifreq's ifr_addr field is a struct sockaddr, which is a fair bit
smaller than struct sockaddr_dl (though big enough to hold an ethernet
address). Existing callers set the sockaddr length to
sizeof(struct sockaddr_dl), which is too large, and causes OOB accesses
when if_findmulti() is used to compare the address with others, or when
if_addmulti() makes a copy.
Fix this without breaking compatibility: copy the user-supplied address
into a sockaddr_dl on the stack, and use the latter for the respective
operation.
Also validate the sockaddr_dl internal length fields, suggested by zlei.
Reported by: Yuxiang Yang, Yizhou Zhao, Ao Wang, Xuewei Feng, Qi Li,
and Ke Xu from Tsinghua University using GLM-5.2 from Z.ai
Reviewed by: zlei, ae, glebius
MFC after: 1 week
Sponsored by: The FreeBSD Foundation
Differential Revision: https://reviews.freebsd.org/D59919
pfi_initialize() registers pfi_rename_ifnet_event() on ifnet_rename_event, but pfi_cleanup() never deregisters it. After "kldunload pf", the next interface rename calls through a stale pointer into the unloaded module: kldload pf kldunload pf ifconfig epair create ifconfig epair0a name foo Reviewed by: kp Approved by: kp (mentor) Fixes: https://cgit.freebsd.org/src/commit/?id=349fcf079ca3 ("net: add ifnet_rename_event EVENTHANDLER(9) for interface renaming") Sponsored by: Rubicon Communications, LLC ("Netgate")
No functional change intended. MFC after: 1 week Sponsored by: The FreeBSD Foundation
There is a chance that the m_pullup() call immediately above invalidated the "ip" pointer, so refresh it as we do with the IGMP header. Otherwise the test in igmp_input_v3_query() for whether the packet is an IGMPv3 general query might use a pointer to a freed mbuf. MFC after: 1 week Sponsored by: The FreeBSD Foundation
PR: https://bugs.freebsd.org/bugzilla/show_bug.cgi?id=298877 MFC after: 1 week Sponsored by: Rubicon Communications, LLC ("Netgate")
The netlink encoding of struct pfr_addr omits pfra_fback, so no per-address feedback reaches userspace. In particular, pfr_get_astats() marks entries of tables without counters with PFR_FB_NOCOUNT, which pfctl uses to skip their counters, so "pfctl -v -T show" prints zero counters for such tables. Add PFR_A_FBACK, emit it from the kernel and decode it in libpfctl. Add a regression test. Reviewed by: kp Approved by: kp (mentor) Fixes: https://cgit.freebsd.org/src/commit/?id=08f54dfca197 ("pf: convert DIOCRGETASTATS to netlink") Sponsored by: Rubicon Communications, LLC ("Netgate")
Removal of opt_route.h left route.c without FIB_ALGO, skipping fib_destroy_rib() in rt_table_destroy(). Explicitly guard it to clean up fib_algo instances and callouts before releasing the rib_lock rmlock, preventing a page fault in softclock and VNET resource leaks. Approved by: pouria Fixes: https://cgit.freebsd.org/src/commit/?id=254b23eb1f5 ("routing: Retire ROUTE_MPATH compile option") Sponsored by: Netflix Differential Revision: https://reviews.freebsd.org/D60158
The purge thread sets PFI_IFLAG_REFS on the interfaces that states
refer to without the rules lock, under which the other flags are
changed. The updates can interleave, so that a "set skip on" is lost,
or outlives its removal, until the next ruleset load.
Use atomic operations to modify the flags. In the purge thread, only
write if the flag is not already set.
Reviewed by: kp
Approved by: kp (mentor)
MFC after: 1 week
Sponsored by: Rubicon Communications, LLC ("Netgate")
Differential Revision: https://reviews.freebsd.org/D60107
pf: free an unparsed rule with pf_krule_free() in pf_handle_addrule() When the PFNL_CMD_ADDRULE message fails to parse, pf_handle_addrule() frees the rule with pf_free_rule(), which asserts the rules and config locks (neither is held) and releases references that pf_ioctl_addrule() has not taken yet. With INVARIANTS this panics on any parse error; without, a rule address parsed as PF_ADDR_TABLE makes pfr_detach_table() dereference NULL. Use pf_krule_free(), as the ioctl paths do. Reviewed by: kp Approved by: kp (mentor) Fixes: https://cgit.freebsd.org/src/commit/?id=e249f5daa41f ("pf: fix memory leak on rule add parse failure") MFC after: 1 week Sponsored by: Rubicon Communications, LLC ("Netgate") Differential Revision: https://reviews.freebsd.org/D60104
pf: leave the epoch to purge unlinked rules pf_purge_thread() calls pf_purge_unlinked_rules() in the network epoch, where sleeping is not allowed, and it takes pf_config_lock, an sx lock. If a rule is being added at the time, the purge thread can sleep on the lock, which panics with INVARIANTS. Reviewed by: kp Approved by: kp (mentor) Fixes: https://cgit.freebsd.org/src/commit/?id=f92d9b1aad73 ("pflow: import from OpenBSD") MFC after: 1 week Sponsored by: Rubicon Communications, LLC ("Netgate") Differential Revision: https://reviews.freebsd.org/D60160
pf: take the rules read lock in pf_handle_getrule() pfctl -sr calls PFNL_CMD_GETRULE once per rule, and pf_handle_getrule() takes the rules write lock each time, so listing a ruleset of N rules stops packet processing N times. Only zeroing the counters (pfctl -z) needs the write lock. Take the read lock otherwise, as DIOCGETRULENV does. Reviewed by: kp Approved by: kp (mentor) Fixes: https://cgit.freebsd.org/src/commit/?id=777a4702c591 ("pf: implement addrule via netlink") MFC after: 1 week Sponsored by: Rubicon Communications, LLC ("Netgate") Differential Revision: https://reviews.freebsd.org/D60161
Stuff in man section 8 (other than networking).
The same variable was used as a counter for an inner and out loop. Add a new one for the inner loop. PR: https://bugs.freebsd.org/bugzilla/show_bug.cgi?id=283934 Reported by: crest at rlwinm.de
This tool supports two commands. The list command outputs a summary of injectable errors supported by the current system. The inject command injects the requested error. Reviewed by: gallatin, imp Sponsored by: Netflix Differential Revision: https://reviews.freebsd.org/D58026
LOG_CONS was OR'd into the facility argument instead of logopt, leaving logopt as 0. The correct call is openlog(ident, LOG_CONS, LOG_AUTH), as shutdown(8) and init(8) already do. PR: https://bugs.freebsd.org/bugzilla/show_bug.cgi?id=296315 Signed-off-by: Ricardo Branco <rbranco@suse.de> Reviewed by: imp, des Pull Request: https://github.com/freebsd/freebsd-src/pull/2300
By default bhyve(8) creates a snapshot socket in "/var/run/bhyve/" (BHYVE_RUN_DIR). As this is a system directory not writable by users, this does not work when bhyve(8) is being started as a non-root user. Address that by allowing to override this directory. In bhyve(8) it is done by setting 'rundir' option with '-o rundir=<path>'. In bhyvectl(8) it is done with '--rundir=<path>'. MFC after: 1 month Reviewed by: bcr (manpages), bnovkov Sponsored by: The FreeBSD Foundation Differential Revision: https://reviews.freebsd.org/D57494
The download protocol calls download_data with FileOffset and BufferLength of 0 first to start the download (no data yet available). Calls it again with BufferLength == 0 and FileOffset the size of the download (again, no data). It then starts calling with BufferLength != 0 and FileOffset == 0 to start the download. The heuristic I used to detect the start was wrong, so we'd allocate the buffer twice. Fix that by being more explicit and not using the heuristic that was bogus. Fixes: https://cgit.freebsd.org/src/commit/?id=afee781523e4 ("loader.efi: Recognize new memdisk=<url> and memcd=<url> options") Sponsored by: Netflix Differential Revision: https://reviews.freebsd.org/D58068
The end address is the final byte in the array, not one byte past the end of the array, so we need to add 1 to get the full length. Fixes: https://cgit.freebsd.org/src/commit/?id=59219fc76a4b ("loader.efi: efiblk_memdisk_preload passes the VirtualDisks to FreeBSD") Sponsored by: Netflix Differential Revision: https://reviews.freebsd.org/D58069
This code is simpler when we spell it the Unix way. Also, add sanity checks to make sure the offset is where we think it is. Fixes: https://cgit.freebsd.org/src/commit/?id=afee781523e4 ("loader.efi: Recognize new memdisk=<url> and memcd=<url> options") Sponsored by: Netflix Differential Revision: https://reviews.freebsd.org/D58070
We have two sets of BIOS loaders: One that lives in stand/i386 and one that lives in stand/userboot. Add knows to turn these on/off, with the default being on. These often aren't needed when creating a minimal UEFI system, so add knobs to turn them off. Given light-weight VMs have created a new use cases for these loaders, there's no plans at all to eliminate them. Sponsored by: Netflix Differential Revision: https://reviews.freebsd.org/D58072
kldxref -m <file> will print the same data that the '-d' flag produces, except restrict the output to one file. This should be the full path to the file, and the directory name to process is omitted. Sponsored by: Netflix Differential Revision: https://reviews.freebsd.org/D57902
Reviewed by: jfree, kib MFC after: 3 weeks Differential Revision: https://reviews.freebsd.org/D58160
We can now decompress .xz compressed memory disks, like FreeBSD-15.1-RELEASE-amd64-disc1.iso.xz Sponsored by: Netflix Differential Revision: https://reviews.freebsd.org/D58073
loader: Add xzfs, like gzipfs but with xz. This is just like gzipfs or bzipfs, except done with the newer xz program. This is off by default for the moment. Sponsored by: Netflix
loader: zstd based filesystem, zstdfs like gzipfs Off by default. Sponsored by: Netflix
loader: Add forgotten xz.c and zstdfs.c These were overlooked when I added compression support. Fixes: https://cgit.freebsd.org/src/commit/?id=86d719ae68aa ("loader: Add xzfs, like gzipfs but with xz.") Fixes: https://cgit.freebsd.org/src/commit/?id=c61ee49cd06a ("loader: zstd based filesystem, zstdfs like gzipfs") Sponsored by: Netflix
zonectl's Report Zones subcommand displays a tabular list of zones. A conventional zone's WP column is displayed as 0xffffffffffffffff , the literal value that the HDD reports. But that's too wide for the column, causing the text to be misaligned. It's also not really meaningful, because the Write Pointer isn't really defined for a Conventional zone. Change it to "-1" to fix the text misalignment. MFC after: 2 weeks Sponsored by: ConnectWise Reviewed by: fuz Differential Revision: https://reviews.freebsd.org/D57512
illumos smatch build is complaining:
pci_nvme_parse_config() warn: 'sc->max_qentries' unsigned <= 0
pci_nvme_parse_config() warn: 'sc->ioslots' unsigned <= 0
Because we are using atoi() to translate string to int, we need
to use int type variable for translation.
Reviewed by: bnovkov
Differential Revision: https://reviews.freebsd.org/D58213
Somehow, I wound up with space indents rather than tab indents, so fix this. Sponsored by: Netflix
For devices like the rtw88, they will show up in `ifconfig -l` as rtw880, rtw881, etc. We want to query the rtw88.0 and rtw88.1 sysctl respectively, not rtw.880. Chances are that there aren't more than 9 wlan devices using the same driver. Use a better heuristic to get the device description. Reviewed by: bz MFC after: 3 days Sponsored by: The FreeBSD Foundation
Modify `blockif_open` to properly release a partially initialized `blockif_ctxt` structure on error. Differential Revision: https://reviews.freebsd.org/D57887 Reviewed by: novel, bnovkov, glebius Tested by: bnovkov MFC after: 2 weeks
Since init become dynamically linked, reroot appeared to be broken because init copies itself into a transient tmpfs mount to continue controlling execution right after the reboot(REROOT) syscall. Because the binary is dynamically linked, it cannot be properly executed. Provide a minimal static binary 'reroot_seed' embedded into the init as byte stream, which performs what the 'init -r' did, namely, the second phase reroot. For the static build of init as part of the /rescue crunch, keep the inline reroot code. Reported and tested by: markj Sponsored by: The FreeBSD Foundation MFC after: 1 week Differential revision: https://reviews.freebsd.org/D58164
pfr_add_table() does not set errno, it returns an error (now).
Read the error code from the return value so we display the correct
error message to the user.
Sponsored by: Rubicon Communications, LLC ("Netgate")
These calls return an error value, they do not set errno. Check their
return values.
Sponsored by: Rubicon Communications, LLC ("Netgate")
Pass such a section to the kernel using modinfo, otherwise link_elf.c won't execute constructors for the file. This is required for KASAN, otherwise redzones for global buffers are not poisoned during boot. Reviewed by: kib MFC after: 2 weeks Sponsored by: The FreeBSD Foundation Differential Revision: https://reviews.freebsd.org/D58244
- Scope local variables properly to each function.
- Quote variables that should be treated as single words.
- Replace `${cmd}; if [ $? -eq 0 ]` with `if ${cmd}` for simplicity.
MFC after: 1 week
Differential Revision: https://reviews.freebsd.org/D57899
Add the option "oemstring" to allow setting the DMI type 11 ("OEM
Strings") SMBIOS structure. These are free-form strings, available for
any purpose, but can be especially useful to pass configuration,
secrets, and credential information into a Linux guest and consumed by
systemd.
MFC after: 1 month
Relnotes: yes
Reviewed by: markj
Differential Revision: https://reviews.freebsd.org/D57516
When writing to a file, call fchmod() to ensure the file mode matches the intended mode, which is 0444. This was already done when replacing an existing file, but not when creating a new file, which meant if the process umask was 077, the resulting certificates and bundle would be unreadable by unprivileged users. MFC after: 1 week Reviewed by: des Differential Revision: https://reviews.freebsd.org/D58304
hwpmc: add regression tests for counting-PMC counter wraparound Exercise a process-mode counting PMC whose accumulated count crosses, or already exceeds, the range of the underlying hardware counter. Before the previous commit, the first context switch after the hardware counter wrapped panicked INVARIANTS kernels with "negative increment" and silently corrupted the accumulated count on other kernels. The tests need a hardware counting event backed by a counter narrower than 64 bits and skip where none is available (hwpmc(4) not loaded, or a VM without a vPMU). Reviewed by: adrian MFC after: 2 weeks Assisted-by: Claude Code (Fable 5) Differential Revision: https://reviews.freebsd.org/D58341
tests/sys/pmc: only build if MK_PMC != no This unbreaks the build when pmc support is explicitly disabled via the aforementioned build knob. MFC after: 10 days Fixes: https://cgit.freebsd.org/src/commit/?id=2cfd82f74 ("hwpmc: add regression tests for ...") Differential Revision: https://reviews.freebsd.org/D58401
vt(4) does not (currently) support changing the video mode. Report that -i mode is not supported rather than printing an empty list. PR: https://bugs.freebsd.org/bugzilla/show_bug.cgi?id=207411 Sponsored by: The FreeBSD Foundation Differential Revision: https://reviews.freebsd.org/D58163
The max_qentries in pci_nvme_softc is uint16_t and too large int may get truncated to invalid value. While there, use local declarations for val. Suggested by: Bill Sommerfeld Reviewed by: chuck Differential Revision: https://reviews.freebsd.org/D58293
Firmware on a test machine applied NX to non-code allocations, which resulted in a fault when jumping to the trampoline. Reviewed by: kib Tested by: Jim Huang Chen <jim.chen.1827@gmail.com> Sponsored by: The FreeBSD Foundation Differential Revision: https://reviews.freebsd.org/D58383
MFC after: 1 week
Commit 9e1db51d4b5fc made nvmecontrol devlist get list of active namespaces from the device instead of iterating through all possible IDs. The problem is that this request is not supported before NVMe 1.1, and in particular by Intel Optane 905P drives. This change reintroduces iteration for devices before NVMe 1.2. Reviewed by: imp Differential Revision: https://reviews.freebsd.org/D58010
Update fwget(8) to download wifi-firmware-mt76-kmod-mt7921, and wifi-firmware-mt76-kmod-mt7925 firmware packages instead of the no longer available mt792x version. Add another PCI vendor to recognize ITTIM IDs for mt7921-based MediaTek cards. (bz reduced the license in the ittim file to an SPDX tag and updated the commit message, given this is only half the work from the review) Sponsored by: The FreeBSD Foundation MFC after: 3 days Differential Revision: https://reviews.freebsd.org/D57242
pmc: console configuration and table rendering for new PMC tools Initializes the terminal rendering code used by the new pmc tools. Then provides a table abstraction for collecting, sorting and rendering tables. It provides pretty printed results with typed fields that print several types used throughout the new PMC tools. By default the fields are formatted in engineering notation. Sponsored by: Netflix Reviewed by: adrian, imp Differential Revision: https://reviews.freebsd.org/D57775
pmc: new pmc log processing framework View is a class for building PMC log processing tools it is designed to work with the new PMC record command that adds a header with additional CPU information. The new framework processes PMC logs about 2.5 times faster and in about half the code as libpmcstat. Sponsored by: Netflix Reviewed by: adrian Differential Revision: https://reviews.freebsd.org/D57776
pmc: pmc record command The record command is designed around the idea of predefined studies. While you can still select individual counters, the predefined studies are meant to enable the best hardware options for a given generation. It implements all of the base studies that I have built so far. Sponsored by: Netflix Reviewed by: adrian Differential Revision: https://reviews.freebsd.org/D57777
pmc: pmc info command Prints the log header including machine, cpu and kernel details along with what counters were selected. Sponsored by: Netflix Reviewed by: adrian, imp Differential Revision: https://reviews.freebsd.org/D57778
pmc: pmc frontend stall analysis based on IBS The frontend command uses AMD IBS frontend events to analyze the major sources of frontend stalls. It displays a table breakind down the major causes of front end stalls. This is a simple demonstration of the tools as you can use the filtering tools to limit the analysis to a subset of the samples including filtering by fetch latencies. Sponsored by: Netflix Reviewed by: adrian, imp Differential Revision: https://reviews.freebsd.org/D57779
pmc: enable the new pmc commands This change hooks everything up to the pmc command and improves the usage to document all functions. There are a couple older commands that are currently broken that I have hidden from the usage, but left in the code for those using it. I won't remove those until we have our replacements upstreamed that depend on the AMD PMC multiplexing patches. Sponsored by: Netflix Reviewed by: adrian, imp Differential Revision: https://reviews.freebsd.org/D57780
route(8): Add prefsrc option in netlink Add prefsrc option that is frequently used on unnumbered interfaces or L3 multi-homed network hosts. This option uses RTA_PREFSRC. Now you can add a static route by specifying the prefsrc option with the loopback IP. PR: https://bugs.freebsd.org/bugzilla/show_bug.cgi?id=285422 Reviewed by: glebius Differential Revision: https://reviews.freebsd.org/D58294
route(8): Add null check for prefsrc option Reported by: Shawn Webb <shawn.webb@hardenedbsd.org>, bms Fixes: https://cgit.freebsd.org/src/commit/?id=dd235f097af4 ("route(8): Add prefsrc option in netlink")
Modify the disk check to allow arbitrary files as the trailing argument instead of requiring a live GEOM disk provider. This enables modifying a boot0 binary file in-place before flashing it to a disk via gpart bootcode, or using it directly as an argument to mkimg's partition specification, as these tools cannot directly adjust the parameters of the boot0 boot manager. Reviewed by: imp, jhb MFC after: 1 week Differential Revision: https://reviews.freebsd.org/D57310
The need for this step is fading, now is mostly used to allow the selection of just the two code partitions in the boot0 boot manager, instead of the default of allowing all four MBR slices (the other two being cfg and data, which cannot boot). Reviewed by: imp MFC after: 1 week Differential Revision: https://reviews.freebsd.org/D57311
The test uses a fail point to inject a decryption error in OCF while sending a ping across the tunnel. The driver should then fail to respond to the ping and increment the input error counter on the interface. Approved by: so Security: FreeBSD-SA-26:52.if_wg Security: CVE-2026-58085 Reviewed by: markj Sponsored by: Chelsio Communications
Reviewed by: kib Fixes: https://cgit.freebsd.org/src/commit/?id=561991144e42 ("Remove Obj_Entry textsize member.") Sponsored by: AFRL, DARPA Differential Revision: https://reviews.freebsd.org/D58522
- When cached response is available, actually use xid from one instead of using its byte-swapped value for BIOS and 1 for UEFI. - If cached response is not available, generate pseudo-random xid, since use of a constant may cause conflicts if two systems are booting same time, and server sends responses as broadcast. - When cached response is available, skip DHCP DISCOVER/OFFER and just send REQUEST to the DHCP server from the cached response. We could skip this phase too and just use the cached response, but we don't know whether firmware requested all of DHCP options we'd like to get. Tested on amd64 Supermicro X11DPI-NT for both BIOS and EFI, with and without cached response packet.
Implement netlink support for gre in ifconfig Differential Revision: https://reviews.freebsd.org/D55366
Transaction ID should persist only between OFFER and the following REQUEST. In all other cases it should change.
This implementation does not cover tunnel addresses. Differential Revision: https://reviews.freebsd.org/D57667
Per RFC1717 section 5.1.3, the option length must be at least three. Processing an undersized option would trigger a large out-of-bounds write. PR: https://bugs.freebsd.org/bugzilla/show_bug.cgi?id=271910 Reported by: Robert Morris Reported by: Décio Brandão (0xDBJ) Reviewed by: emaste MFC after: 1 week Sponsored by: The FreeBSD Foundation Differential Revision: https://reviews.freebsd.org/D58554
Some drivers do this routinely, e.g., FreeBSD's tpm20 does this every time it sends a command in tpmcrb_transmit(). This causes the console to fill up with messages. Instead, only print a warning if the cancel bit is set to one. Reviewed by: corvink MFC after: 2 weeks Differential Revision: https://reviews.freebsd.org/D52425
Commit 505222d35fea removed a batch of code that this goto used to skip around. Reviewed by: olce, kib, markj Fixes: https://cgit.freebsd.org/src/commit/?id=505222d35fea ("Implement the long-awaited module->file cache database. A userland tool (kldxref(8)) keeps a cache of what modules and versions are inside what .ko files. I have tested this on both Alpha and i386.") Differential Revision: https://reviews.freebsd.org/D58539
Previously all the 'goto out' statements after the image was loaded into memory returned success rather than an error. This is despite comments indicating some of these conditions were in fact errors, and some of these error conditions (such as missing PT_DYNAMIC) are treated as errors in the kernel linker. In addition, when failing to looking up the symbols for the linker set, those cases returned failure leaking memory (though it's clear from the original code from commit ca49b3342d1e that only the second failure was intended to be an actual error). To avoid more confusion, move the assignment of `ret` to just before the `out` label so that `goto out` always returns an error. This is a more consistent pattern with other code in the tree that tends to use labels for the error case. Restructure some other code to avoid a few bogus errors. Specifically, a symbol table is not required so don't treat lack of a symbol table as an error. Also, if the start symbol for the module metadata linker set is not found, don't treat that as an error either. Reviewed by: kib Differential Revision: https://reviews.freebsd.org/D58540
All sorts of places in the ELF loading code assume that filesz <= memsz, so check that explicitly up front. The kernel already performs this check for the PT_LOAD segments in the main binary and rtld in imgact_elf.c. Reviewed by: jrtc27, kib Differential Revision: https://reviews.freebsd.org/D58541
Reviewed by: jrtc27, kib Differential Revision: https://reviews.freebsd.org/D58543
Pass a single module name to load_kld for kbdmux and vkbd, allowing bthidd_prestart to load both modules successfully. Fixes: https://cgit.freebsd.org/src/commit/?id=cfe1962a1925 (rc: Fix improper use of load_kld) MFC after: 3 days Sponsored by: The FreeBSD Foundation
Allowing nuageinit user scripts to run before these makes it possible to customize official BASIC-CI and BASIC-CLOUDINIT FreeBSD images. This was requested by KDE for their CI. Approved by: cperciva Pull-Request: https://ron-dev.freebsd.org/FreeBSD/src/pulls/60
Each byte of the address is represented by a pair of characters, so we should be multiplying len by 2 when figuring out how much buffer space we have. Previously, a sufficiently large option could cause an overflow of the global "result" buffer. Reported by: Joshua Rogers <joshua@joshua.hu> Tested by: Décio Brandão (0xDBJ) MFC after: 3 days Sponsored by: The FreeBSD Foundation Differential Revision: https://reviews.freebsd.org/D58555
This is in lieu of a full Kyua/ATF regression test, as this is an optional feature that was beyond the scope of IETF's normative references for IPv6 multicast; support has been strictly on a best-effort basis. Two new commands are added to mtest(8): u mcast-addr ifname - join IPv4-mapped group on IPv6 socket v mcast-addr ifname - leave IPv4-mapped group on IPv6 socket Add an internal helper function __in6_v4_to_v4mapped() to perform the converse of the IN6_IS_ADDR_V4MAPPED() check to support this use case. Whilst __in6_v4_to_v4mapped() returns its first argument as a convenience, avoid the temptation to dereference a pointer to that which we already hold. Strictly the use of sockunion_t within mtest(8) more generally is a form of controlled type punning (aliasing). Use a temporary as we overwrite contents of su; the resultant write would overlap memory locations. PR: https://bugs.freebsd.org/bugzilla/show_bug.cgi?id=https://bugs.freebsd.org/bugzilla/show_bug.cgi?id=193246 Differential Revision: https://reviews.freebsd.org/D58589
Fixes: https://cgit.freebsd.org/src/commit/?id=7e2f38311e62 ("rtld-elf/rtld.c: apply clang-format") Sponsored by: Innovate UK
Populate stand/libsa/bootp.c's bootp_response global from the UEFI PXE Base Code Protocol's cached DhcpAck, so bootp() can enter RFC 2131 INIT-REBOOT and skip DISCOVER/OFFER instead of running a fresh DHCP transaction after the firmware has already done one.
It should give DHCP servers more information for proper responses.
During flag inconsistency report, we handle rai->rai_otherflg as a bool, but the value is 0x40. Make it a simple number comparison. PR: https://bugs.freebsd.org/bugzilla/show_bug.cgi?id=295995 Reviewed by: markj, Faraz Vahedi <kfv@kfv.io> MFC after: 3 days Differential Revision: https://reviews.freebsd.org/D58672
niov2 returns the number of entries in the iovec starting at offset "offset". Here we are unconditionally setting it to 1, which of course isn't right. Fixes: https://cgit.freebsd.org/src/commit/?id=a28cf86c4171 ("bhyve/virtio: Rework iovec handling functions for efficiency and clarity") Reported by: Claude and Ada Logics Reviewed by: Hans Rosenfeld <rosenfeld@grumpf.hope-2000.org> MFC after: 1 week Sponsored by: The FreeBSD Foundation Differential Revision: https://reviews.freebsd.org/D58625
bsdinstall: add a hardening knob for unprivileged kenv access It makes sense. Reviewed by: zleei Differental Revision: https://reviews.freebsd.org/D57755
kern: fix oversight in security.bsd.unprivileged_kenv_read It was intended that one could close the hole back in loader, but the sysctl was actually not marked TUNABLE. The hardening menu option thus did nothing, because we wouldn't read the value from kenv. Reported by: markj Fixes: https://cgit.freebsd.org/src/commit/?id=6e81fbf5833d ("bsdinstall: add a hardening knob [...]") Fixes: https://cgit.freebsd.org/src/commit/?id=4fd518fcb2bb ("kern: add a security knob to disable [...]")
We can probaby consider these kernel bugs, in which case asserting is not the most helpful thing we can do. Let's emit the necessary details to stderr and exit non-zero to aid debugging these without completely blocking the ability to export all of the well-formed metrics. Reviewed by: rew Differential Revision: https://reviews.freebsd.org/D57983
PeriphID and CellID values are determined by macros which take an index. They currently receive a bus offset which has a stride of 4 bytes. This causes the ID1-3 registers to report incorrect values. Scale the offset before passing it to the macro to fix this. Tested with kvm-unit-tests/arm/pl031. Signed-off-by: Kajetan Puchalski <kajetan.puchalski@arm.com> Reviewed by: jrtc27 Fixes: https://cgit.freebsd.org/src/commit/?id=014d7082a239 ("bhyve: Implement a PL031 RTC on arm64") MFC after: 1 week Pull Request: https://github.com/freebsd/freebsd-src/pull/2358 Closes: https://github.com/freebsd/freebsd-src/pull/2358
Add -L to query the generic packed-nvlist IOV_GET_STATUS interface. Report PF enable state and configured and total VF counts. For each VF, print its PCI address, newbus attachment, bound driver, and ppt state. Retry size negotiation if the topology changes between ioctls and reject malformed or incompatible status records. Keep NIC-specific operational state in ifconfig -v; iovctl owns the device-neutral PCI topology and applies to any SR-IOV device class. Relnotes: yes
- Prefix all structs with the struct keyword to avoid collisions between the types and variables with the same "name". - Use `_` suffixed variables in initializers to distinguish input parameters from public members [1]. Resolve some trailing whitespace issues while here. NOTE: this doesn't resolve the -pedantic issue reported by g++ with `pmchdr_cpuidinfo::cpuid` about the field being a flexible array in an otherwise empty struct. 1. I generally do this the other way around, i.e., suffix private/protected members with `_`, but these are public members in structs and I don't want to introduce a lot of churn in calling code. Reported by: g++14 with FreeBSD CI (powerpc64 tinderbox) Fixes: https://cgit.freebsd.org/src/commit/?id=ce6ab51f ("pmc: enable the new pmc commands")
This mutes a number of complains from g++ about needing specific headers for functionality related to C strings and other function prototypes. Reported by: g++ 14
I need to do more work before references can be accepted in other sections of the code. This was an unnecessary drive-by change that was not tested in `make universe`. Reported by: CI Fixes: https://cgit.freebsd.org/src/commit/?id=fd809148 ("pmc(8): resolve -Wshadow issues")
nuageinit: adopt cloud-init disable_root semantics disable_root now restricts root's authorized_keys instead of setting PermitRootLogin. Reported by: np@
nuageinit: fix ssh_pwauth string handling Treat "no"/"unchanged" correctly instead of any non-nil value as yes.
nuageinit: accept lock_passwd for users Alias cloud-init lock_passwd key alongside locked.
nuageinit: support allow_public_ssh_keys Skip importing datasource public keys when set to false.
Reviewed by: markj Differential Revision: https://reviews.freebsd.org/D58636
bhyve: namescope virtio_msix to virtio.msix The bhyve_config(5) variable `virtio_msix` is namescoped to `virtio.msix`. Configurations that have the old variable will automatically be mapped to the new one, with a warning message printed out. Relnotes: yes Reviewed by: ziaee, markj Differential Revision: https://reviews.freebsd.org/D58390
fbsdrun_virtio_msix(): update virtio_msix to virtio.msix Fixes: https://cgit.freebsd.org/src/commit/?id=2d985d577d79 ("bhyve: namescope virtio_msix to virtio.msix") Reviewed by: markj Differential Revision: https://reviews.freebsd.org/D58787
Allow "legacy" alongside "none" as a valid value for the ZFS mountpoint property, matching zfsprops(7). Reviewed by: imp, markj MFC after: 1 week Differential Revision: https://reviews.freebsd.org/D58781
nuageinit_user_data_script references 'firstboot_freebsd_update' and 'firstboot_pkg_upgrade', which are from Ports. In a default base system test without sysutils/firstboot-freebsd-update and sysutils/firstboot-pkg-upgrade, rcorder will warn on "unknown provisions" to stderr, but is otherwise harmless. Reviewed by: arrowd Fixes: https://cgit.freebsd.org/src/commit/?id=16e47f317c4ce2be5fed530bf8a9af9f9bf55364 MFC after: 3 days Sponsored by: The FreeBSD Foundation
Co-authored-by: Michael Osipov <michaelo@FreeBSD.org> PR: https://bugs.freebsd.org/bugzilla/show_bug.cgi?id=280487 Reviewed by: kevans, michaelo MFC after: 1 week Differential Revision: https://reviews.freebsd.org/D46313
rtld has always been built PIC since commit 7ca8e6a67068e8357e251bd3ea86253c8a751d59. The stale #ifdef might confuse a reader by thinking rtld can be built as non-PIC. Reviewed by: kib Sponsored by: AFRL, DARPA Differential Revision: https://reviews.freebsd.org/D58623
The name of the script and the name used internally for rc.conf differ, as such the hardcoded disabling of service jails for the didn't work. Fix by using the correct name. Fixes: https://cgit.freebsd.org/src/commit/?id=f99f0ee14e3af rc.d: add a service jails config to all base system services
We have ports and basesystem services, where the internal name and the filename differ. While the documentation recommends to keep them in sync, the reality is different. For service jails use the basename of the service filename. Fixes: https://cgit.freebsd.org/src/commit/?id=2efbd48 rc: add service jails framework Suggested by: joneum MFC after: 1 week MFC to: stable/15
nfsd: Update the rc.d script for RDMA for the nfsd service Commit 7144a1d58c5c added the hooks for the nfsrdma.ko module. Once loaded, this module adds RDMA support to the nfsd. This patch adds a few lines to /etc/rc.d/nfsd, so that nfs_server_rdma_enable="YES" in your /etc/rc.conf will load nfsrdma.ko, so that RDMA service is enabled. It also supports nfs_server_rdma_listen="port#" so that the default of 20490 can be overridden in /etc/rc.conf. At this available as time, the nfsrdma.ko module is an unofficial port, since it was developed by Vinicius Ferrao <ferrao@versatushpc.com.br> using generative AI. As soon as it is available, it will be announced on freebsd-current@freebsd.org. Suggested by: Vinicius Ferrao <versatushpc.com.br> MFC after: 1 month
rc.conf: Fix the default NFS-over-RDMA port number The default for nfs_server_rdma_listen transposed two digits: 20490 instead of 20049, the IANA-assigned port for NFS-over-RDMA. Fixes: https://cgit.freebsd.org/src/commit/?id=471e14267bea ("nfsd: Update the rc.d script for RDMA for the nfsd service") MFC after: 1 month Sponsored by: VersatusHPC Pull Request: #2371 Signed-off-by: Vinícius Ferrão <ferrao@versatushpc.com.br>
Currently, sending SIGTERM to the bhyve process triggers ACPI poweroff for a VM. However, when running bhyve in monitor mode (-M), there are two processes: the monitor process and the actual VM process. Sending SIGTERM to the VM process works as before -- it powers off the VM. But sending SIGTERM to the monitor process just kills the monitor process, leaving the stale VM process running. Fix that by creating a pipe between these two processes. The child process uses the pipe to detect when the monitor goes away, and exits automatically. MFC after: 2 weeks Reviewed by: markj Sponsored by: The FreeBSD Foundation Differential Revision: https://reviews.freebsd.org/D58788
bsdconfig: add datetime module for live system clock Provide bsdconfig datetime (menu), date, and time to set the running system clock via dialog(1)/bsddialog(1) --calendar and --timebox with mustberoot. Unlike bsdinstall's time helper, this operates on the live system rather than a CHROOT-staged install target, and calls adjkerntz(8) after a successful change so the CMOS stays in sync. Reviewed by: bcr Differential Revision: https://reviews.freebsd.org/D58487
mtree: usr: add missing bsdconfig datetime directories 8cfe06ee4415 installs into 085.datetime and share/bsdconfig/datetime, but those paths were not in BSD.usr.dist. In-place installworld then fails when install(1) cannot create the destination. Fixes: https://cgit.freebsd.org/src/commit/?id=8cfe06ee4415 ("bsdconfig: add datetime module for live system clock")
Merge implementation of "traceroute -e" to traceroute6 for TCP/UDP/SCTP. MFC-after: 2 weeks
On ZAC drives, conventional zones conventionally report a write pointer LBA of 0xffffffffffff. This field is 48 bits wide, unlike ZBC's 64 bits. Recognize both ZAC and ZBC all-ones behaviour in the WRITE POINTER LBA field to indicate non-valid information. Tested by: fuz Discussed with: fuz, asomers, ken Fixes: https://cgit.freebsd.org/src/commit/?id=4735ef6196bc ("zonectl: display conventional zones better") MFC after: 2 weeks Sponsored by: Google Summer Of Code 2026 Reviewed by: asomers Pull Request: https://github.com/freebsd/freebsd-src/pull/2345
No functional change intended. Submitted by: Pedro Giffuni <pfg@freebsd.org> MFC-after: 1 week
The loader's ZFS implementation never set st_dev or st_ino in zfs_dnode_stat(). With an uninitialized struct stat, veriexec's device comparison in lib/libsecureboot/veopen.c read stack garbage and skipped the matching manifest entry, failing with a spurious "no entry" on ZFS root under UEFI Secure Boot. Rather than zeroing the device (which would break veriexec's ability to tell apart the same path on different datasets), populate st_dev and st_ino with the same intrinsic identifiers the kernel uses: - st_dev = the dataset's ds_fsid_guid (as the kernel does via dmu_objset_fsid_guid()/dsl_dataset_fsid_guid()), already read in zfs_mount_dataset() and now propagated through struct zfsmount. - st_ino = the object number resolved in zfs_lookup(), propagated through struct file (the loader's equivalent of the kernel's z_id). dev_t and ino_t are 64-bit on FreeBSD, so both are assigned directly with no hashing. A memset() at the top of zfs_dnode_stat() zeroes the remaining fields so they no longer hold stack garbage. Tested on 16.0-CURRENT (amd64), ZFS-on-GELI root: rebuilt and re-signed the EFI loader; the system boots under UEFI Secure Boot with mac_veriexec active. (boot1 segment was modified by kevans) PR: https://bugs.freebsd.org/bugzilla/show_bug.cgi?id=295935 Sponsored by: Defenso Reviewed-by: kevans Pull-Request: https://github.com/freebsd/freebsd-src/pull/2271
Changes for jng 2.0 -> 9.0 include:
+ Use ng_bridge(4) uplink hooks on ng_ether(4) lower so the host
mapping table stays small (first hook is uplink; unknown unicast
goes only to uplink)
+ Add `jng pin [-h] {-a | NAME ...}' to plant eiface MACs with
NGM_BRIDGE_MOVE_HOST and raise maxStaleness so they do not expire
+ Remove experimental NG_TYPE=iface / ng_tcpmss(4); ng_iface(4)
cannot work with ng_bridge(4)
+ Add -v
+ SPDX-License-Identifier: BSD-2-Clause; bump copyright to 2026
See D58902 for the ng_bridge(4) data-path MOVE_HOST fix.
MFC after: 1 week
Reviewed by: kfv, jlduran
Differential Revision: https://reviews.freebsd.org/D58903
Update examples and comments to return eifaces in exec.prestop with ifconfig -vnet (before jng shutdown in poststop). Reported by: jlduran PR: https://bugs.freebsd.org/bugzilla/show_bug.cgi?id=268397 MFC after: 1 week Reviewed by: kfv, jlduran Differential Revision: https://reviews.freebsd.org/D58939
Use jail.conf(5) $name in the jail.conf examples so the jail name need only be set on the stanza. Keep host.hostname as xxx.yyy; a jail name is not a DNS label. Leave the rc.conf excerpt as xxx. Suggested by: jlduran MFC after: 1 week Reviewed by: kfv, jlduran Differential Revision: https://reviews.freebsd.org/D59032
This pacifies shadow warnings from GCC:
usr.sbin/pmc/cmd_pmc_record.cc: In constructor 'pmc_config::pmc_config(const std::__1::string&, uint64_t, cpuset_t)':
usr.sbin/pmc/cmd_pmc_record.cc:96:71: error: declaration of 'cpumask' shadows a member of 'pmc_config' [-Werror=shadow]
96 | pmc_config(const std::string &event, uint64_t count, cpuset_t cpumask)
| ~~~~~~~~~^~~~~~~
usr.sbin/pmc/cmd_pmc_record.cc:90:25: note: shadowed declaration is here
90 | cpuset_t cpumask;
| ^~~~~~~
usr.sbin/pmc/cmd_pmc_record.cc:96:55: error: declaration of 'count' shadows a member of 'pmc_config' [-Werror=shadow]
96 | pmc_config(const std::string &event, uint64_t count, cpuset_t cpumask)
| ~~~~~~~~~^~~~~
usr.sbin/pmc/cmd_pmc_record.cc:89:25: note: shadowed declaration is here
89 | uint64_t count;
| ^~~~~
usr.sbin/pmc/cmd_pmc_record.cc:96:39: error: declaration of 'event' shadows a member of 'pmc_config' [-Werror=shadow]
96 | pmc_config(const std::string &event, uint64_t count, cpuset_t cpumask)
| ~~~~~~~~~~~~~~~~~~~^~~~~
usr.sbin/pmc/cmd_pmc_record.cc:88:25: note: shadowed declaration is here
88 | std::string event;
| ^~~~~
Reported by: GCC 15
Fixes: https://cgit.freebsd.org/src/commit/?id=a79a051e7d16 ("pmc: pmc record command")
bhyve emulates the guest PCI Command register so BAR sizing does not disable physical decoding. However, PCIe Device Control was passed through. A guest VFIO reset therefore performed a physical FLR, which cleared physical Command, while the guest restored only its emulated copy. The device remained assigned with bus mastering disabled and could not fetch DMA descriptors. Intercept guest FLR writes and issue a PPT-managed reset. Stop all vCPUs, verify ownership, quiesce the function, perform only an FLR, and restore the host-owned PCI configuration, decode, and bus-master state. Keep the IOMMU domain in place. bhyve removes guest BAR mappings before this ioctl; a later guest MEMEN write recreates them. Never escalate a guest FLR to a power reset. Reset the guest-owned Command, MSI, MSI-X, MSI-X table, INTx, and MRRS state. PCIe 6.2 section 6.6.2 explicitly preserves MPS across FLR. Virtualize MPS, MRRS, and Completion Timeout. Keep physical MPS and completion-timeout policy host-owned, and apply physical MRRS with MPS as its floor. Keep Phantom Functions Enable host-owned because it changes requester identities visible to the IOMMU. Serialize guest configuration transactions per function and gate trapped and direct BAR access across the reset. Handle byte, word, dword, and overlapping Device Control accesses. A guest FLR can sleep for at least 100 ms. Reserve the target function while dropping the global PPT lock so a guest cannot delay PPT lifecycle operations for other VMs. Operations on the target wait for its reset while other functions and VMs can proceed. Check pcie_flr_supported() before destructive preparation so PPT applies the generic PCI quirk policy. This includes VFs such as the 82599 which implement FLR without advertising it. Validated with two E610 VFs in a Linux 7.0 guest using VFIO no-IOMMU and DPDK testpmd with two queues per VF. Byte, word, dword, and overlapping FLR writes, 32 alternating resets, DPDK traffic, and ixgbevf reattachment all completed while the sibling VF and host PCIe remained healthy. This fixes Linux VFIO no-IOMMU with DPDK PMDs. Reviewed by: markj MFC after: 2 weeks Sponsored by: BBOX.io
The passthrough Command register is emulated, but PMCSR writes were sent directly to the physical function. A guest D3hot-to-D0 transition can perform an internal reset and clear physical Command while its emulated copy remains enabled. Cache the Power Management capability and keep the physical D-state host-owned. Emulate the guest D-state and advertise No_Soft_Reset so the guest is not promised a function reset by a virtual power cycle. Restore the assignment-time virtual state after a managed FLR. Reviewed by: markj MFC after: 2 weeks Sponsored by: BBOX.io
This allows specifying custom vdev paths (such as GPT labels like /dev/gpt/...) when creating ZFS filesystem images via makefs(8), rather than defaulting to /dev/null. Reviewed by: markj MFC after: 1 week Differential Revision: https://reviews.freebsd.org/D59031
The previous check did not work on architectures where `char` is
unsigned as noted by GCC on aarch64:
usr.sbin/pmc/view.cc: In member function 'void pmcview::loadsymboltable(image*, Elf*, Elf_Scn*, GElf_Shdr*)':
usr.sbin/pmc/view.cc:627:38: error: comparison is always false due to limited range of data type [-Werror=type-limits]
627 | if (fname[i] < 0) {
| ~~~~~~~~~^~~
Reported by: GCC 15
Add $Version, -h/-v, and SPDX-License-Identifier: BSD-2-Clause. Drop the long-form license and bump the copyright to 2026. MFC after: 1 week
Interface names have no limitations on the allowed character set, only a length restriction. Widen the allowed character set for interface names to include any printable character. PR: https://bugs.freebsd.org/bugzilla/show_bug.cgi?id=290916 Reviewed by: dteske MFC after: 1 week Differential Revision: https://reviews.freebsd.org/D53865
Before the preamble script is sourced, initialize BSDINSTALL_LOG with the file in $debugFile, if it is not already set. Without this change, the bsdinstall script would assume the preamble set BSDINSTALL_LOG to empty and the comparison with $debugFile will fail, causing the log to be re-initialized to /dev/null. Reviewed by: dteske MFC after: 1 week Differential Revision: https://reviews.freebsd.org/D25343
PermitRootLogin is "no" by default and that stopped root from logging in even though disable_root was set to false during initialization. Reviewed by: bapt Sponsored by: Chelsio Communications Differential Revision: https://reviews.freebsd.org/D59101
The emulated HDA controller passed values taken from guest registers and from guest memory straight into assert(), so a guest could abort bhyve with values the emulation did not expect. Reject them instead. In case the guest asked to start something and it failed, clear the corresponding run/enable bit. PR: https://bugs.freebsd.org/bugzilla/show_bug.cgi?id=256379, https://bugs.freebsd.org/bugzilla/show_bug.cgi?id=256381, https://bugs.freebsd.org/bugzilla/show_bug.cgi?id=256382, https://bugs.freebsd.org/bugzilla/show_bug.cgi?id=256383, https://bugs.freebsd.org/bugzilla/show_bug.cgi?id=256384, https://bugs.freebsd.org/bugzilla/show_bug.cgi?id=256385, https://bugs.freebsd.org/bugzilla/show_bug.cgi?id=256386, https://bugs.freebsd.org/bugzilla/show_bug.cgi?id=256498 Sponsored by: The FreeBSD Foundation MFC after: 2 weeks Reviewed by: bnovkov, jhb Differential Revision: https://reviews.freebsd.org/D59082
Move the code for the snapshotting IPC thread into a separate file and define macros for adding new IPC commands. No functional change intended. Reviewed by: rew Differential Revision: https://reviews.freebsd.org/D54650
Move the nvlist-based bhyve IPC code into a separate function. No functional change intended. Reviewed by: rew Differential Revision: https://reviews.freebsd.org/D54652
Reported by: Reo Shiseki MFC after: 3 days Sponsored by: The FreeBSD Foundation Differential Revision: https://reviews.freebsd.org/D59054
PR: https://bugs.freebsd.org/bugzilla/show_bug.cgi?id=297839 MFC after: 1 week Obtained from: OpenBSD, henning <henning@openbsd.org>, 5b6657d4d8 Sponsored by: Rubicon Communications, LLC ("Netgate")
Manually specified eui64 value gets converted to big endian twice: first using htobe64() and then using be64enc(). On little-endian hosts that results in a little-endian value instead of a big-endian. Fix by removing htobe64() for a user submitted value. Fixes: https://cgit.freebsd.org/src/commit/?id=409a80e5a434 ("bhyve: Create EUI64 for NVMe namespaces") Reviewed by: chuck Relnotes: yes Sponsored by: The FreeBSD Foundation MFC after: 3 weeks Differential Revision: https://reviews.freebsd.org/D59080
bhyve: Return void from pci_emul_alloc_bar This function never fails. Reviewed by: bnovkov, chuck, markj Differential Revision: https://reviews.freebsd.org/D58579
bhyve: Don't set the prefetch flag for large 64-bit memory BARs The only device model that can create a large 64-bit memory BAR is the passthru device model, and that device model reuses the lobits of the existing BAR explicitly. Fixes: https://cgit.freebsd.org/src/commit/?id=e87a6f3ef284 ("bhyve: use physical lobits for BARs of passthru devices")
bhyve: Refactor initial PCI BAR setup Fully initialize BARs with an address of 0 in pci_emul_alloc_bar() instead of deferring some of that initialization to pci_emul_assign_bar(). Now, the latter is only used to allocate an initial address range for PCI BARs. Note that this means that the pci_passthru model now overrides the initial lobits after they are set removing the need for a workaround in pci_emul_assign_bar(). Reviewed by: bnovkov Differential Revision: https://reviews.freebsd.org/D58893
bhyve: Tidy lobits handling in pci_passthru - The lobits field in the "physical" BAR settings is never used, so don't bother setting it. - Expand the comment explaining why the existing lobits are preserved (namely, to preserve the prefetch flag on memory BARs). Reviewed by: bnovkov Differential Revision: https://reviews.freebsd.org/D58894
Read interface groups from IFLAF_GROUP netlink attribute. Reviewed by: glebius Differential Revision: https://reviews.freebsd.org/D58644
Use better diagnostic messages when unit numbers are wrong. Reviewed by: arrowd@, christos@, kevans@ Approved by: christos@ Differential Revision: https://reviews.freebsd.org/D56845
A service running under ${name}_user was signalled from the host as that
user, which the parent of a jail may no longer do: since 8a5ceebece03 an
unprivileged process would need allow.unprivileged_parent_tampering.
Stop and reload therefore failed and left both the service and its jail
running.
MFC after: 1 week
MFC to: stable/15
rc.subr: svcj - let svcj_all_enable enable service jails Fix the logic for svcj_all_enable. Fixes: https://cgit.freebsd.org/src/commit/?id=2efbd480f1d3 rc: add service jails framework MFC after: 1 week MFC to: stable/15
rc.subr: svcj - remove the service jail when the service is not running
A service whose tracked process had died while another process of its own
kept the jail alive, therefore left svcj-${name} behind, and the next start
would fail.
Fixes: https://cgit.freebsd.org/src/commit/?id=2efbd480f1d3 rc: add service jails framework
MFC after: 1 week
MFC to: stable/15
Assisted-by: Claude Code (Opus 5)
rc.subr: svcj - run a service's own restart and status methods in its jail A script that defines non-default restart_cmd or status_cmd should execute them in the service jail. Where there is no jail to enter, restart starts the service instead of failing. Fixes: https://cgit.freebsd.org/src/commit/?id=2efbd480f1d3 rc: add service jails framework MFC after: 1 week MFC to: stable/15 Assisted-by: Claude Code (Opus 5)
rc.subr: svcj - add a setaudit option
setaudit(8) is prefixed to the command inside the jail when
${name}_audit_user is set, and needs allow.setaudit.
This is not added automatically when ${name}_audit_user is set, this
needs an administrative setting of the options on purpose.
MFC after: 1 week
MFC to: stable/15
Twenty cases over where each rc option and each method executes for a jailed service, the jail's lifetime, and the svcj option handling. Each case drives the service inside a chroot built in its ATF work directory. MFC after: 1 week MFC to: stable/15 Assisted-by: Claude Code (Opus 5)
This patch adds a new nfs_client_rdma_enable variable to /etc/rc.d/nfsclient to enable the client side of NFS over RDMA. The client side of NFS over RDMA requires the nfsclrdma.ko module, which is still under test/review. I wanted to get the "glue" into main so that others could test the module more easily. Avaliability of the module will be announced on freebsd-current@ soon. It should not affect non-RDMA operation. I've specified a long MFC, since the module still requires extensive testing and, hopefully, a review. MFC after: 3 months
In a WITHOUT_INET6 build the only assignment to netid2 is compiled out and the non-INET6 arm returns early, so netid2 is unconditionally NULL and the rpcb_set() call guarded by it is dead code. clang does not prove it dead and reports nbuf2 as uninitialized where it is passed as a const pointer, No functional change. MFC after: 1 week Reported by: clang (-Wuninitialized-const-pointer) Suggested by: dim Approved by: ngie (co-mentor) Reviewed by: ngie Differential Revision: https://reviews.freebsd.org/D59277
The strtoul function sets errno on error, but does not clear it on success; when using strtoul and checking errno (as one should) for ERANGE / EINVAL afterwards, it's important to zero errno first. While here, remove a dead store. Reviewed by: phk Fixes: https://cgit.freebsd.org/src/commit/?id=4fe8c1b67be0 ("Add error and range checking ... ") MFC after: 1 week Differential Revision: https://reviews.freebsd.org/D59270
Pandaboard (sys/arm/ti/omap4) is removed due to lack of HW. Remove the pandaboard config file for nanobsd aswell. Approved by: imp, jlduran, manu(mentor) Diffrential revision: https://reviews.freebsd.org/D54319
Signed-off-by: Aryan Arora <aryanarora.w1@gmail.com> Reviewed by: imp Pull Request: https://github.com/freebsd/freebsd-src/pull/2359
"Qualcomm Atheros Communications Dell Wireless 1802 Bluetooth 4.0 LE" (0cf3:e006) is a wifi-bluetooth combo. The bluetooth chip is confirmed to be AR3012 compatible. That's what the Linux ath3k driver loads as well. It has been tested with the firmware files from the https://git.kernel.org/pub/scm/linux/kernel/git/firmware/linux-firmware.git repository as the comms/ath3k-firmware port appears to be discontinued. Signed-off-by: Robin Haberkorn <rhaberkorn@fmsbw.de> Reviewed by: imp Pull Request: https://github.com/freebsd/freebsd-src/pull/2280
The default on my laptop is annoyingly bright, and this is a useful feature to mitigate that. The backlight script is largely a copy of the mixer service which provides the same value for mixers, but this one is specifically dependant on kld to allow DRM drivers a chance to attach. Note that it's off by default to avoid interference with DEs, and document the capability in backlight(8). Set backlight_enable=YES in rc.conf(5) to enable save/restore. Relnotes: maybe Reviewed by: bapt, ivy, manu, ziaee Differential Revision: https://reviews.freebsd.org/D59296
On SIGHUP reload, closelogfiles() frees each F_PIPE filed even when its pipe process is still running. close_filed() sets f_type to F_UNUSED before the check, so the condition f_type != F_PIPE is always true and the filed is freed while its process descriptor is still on the dead queue and registered in the kqueue. When the child later exits, the NOTE_EXIT handler dereferences the freed filed (use-after-free) and never closes the process descriptor, leaving the pipe child as a persistent zombie. Capture whether the filed is a pipe with an active process descriptor before calling close_filed(), and defer the free in that case so the NOTE_EXIT handler can reap the child and free the filed. Reviewed by: markj Fixes: https://cgit.freebsd.org/src/commit/?id=95381c0139d6 (syslogd: Use process descriptors) Differential Revision: https://reviews.freebsd.org/D59319
While here, use caph_rights_limit(), as syslogd already uses caph_enter(). PR: https://bugs.freebsd.org/bugzilla/show_bug.cgi?id=298104 Reported by: mi Fixes: https://cgit.freebsd.org/src/commit/?id=24816abb8740 ("syslogd: Limit rights on procdescs") MFC after: 3 days
checklost() scans for lost cluster chains and attempts to repair each one, first by reconnecting it to LOST.DIR, and falling back to clearing it if reconnection fails. However, checklost() incorrectly updates the modification status flags (mod), which checkfilesys() relies on to determine whether to write back changes and what exit status to return. The current code have three issues: 1. A reconnect() failure immediately sets FSERROR in mod via "mod |= ret = reconnect(...)". If reconnect() failed (e.g., because LOST.DIR is missing, or full) but the fallback clear operation succeeds, clearchain() frees the chain and sets FSFATMOD. However, the leftover FSERROR remains in mod: checkfilesys() skips marking the file system clean and exits with status 8, even though the file system was fully repaired and a subsequent run finds nothing left to do. This can happen with "fsck_msdosfs -y" on a volume without LOST.DIR. 2. The same assignment overwrites checkchain()'s return value before it can be recorded in mod. When checkchain() truncates a chain (e.g., one whose tail points to a free cluster), its FSFATMOD status is lost, causing checkfilesys() to skip updating the FAT and discard the truncation. With LOST.DIR present so that reconnect() succeeds, "fsck_msdosfs -y" reports "Truncate? yes" and "FILE SYSTEM WAS MODIFIED", exits 0, but leaves the identical damage on disk to be found again on every subsequent run. Similarly, an FSFATAL return value from checkchain() is dropped, defeating the "if (mod & FSFATAL) break" guard that follows. 3. When checkchain() returns FSERROR and clearing the chain is declined, no error status is recorded in mod, causing fsck_msdosfs to report a clean exit despite leaving un-repaired damage. Fix these issues by: - Merge checkchain()'s status into mod before calling reconnect(), and skip reconnect() if checkchain() returned a fatal error. - Defer recording FSERROR from a failed reconnect() until after the Clear fallback attempt, setting FSERROR only if the chain remains unhandled. MFC after: 1 week Pull Request: https://github.com/freebsd/freebsd-src/pull/2251
Add an ATF test suite covering Phase 3 ("Checking for Lost Files")
error accounting. Test images are created using newfs_msdos(8),
and lost cluster chains are injected directly into FAT copies at
offsets derived from the BPB. The LOST.DIR directory required by
reconnect() is constructed similarly: a root directory entry with
ATTR_DIRECTORY set and its first cluster pointing to a zero-filled
cluster containing "." and ".." entries.
The lost_chain_cleared and corrupted_lost_chain_reconnected test
cases provide regression coverage for the preceding commit:
- lost_chain_cleared verifies that clearing a lost chain (the fallback
taken when LOST.DIR is absent) exits with status 0 rather than 8
(unrecovered error).
- corrupted_lost_chain_reconnected verifies that FAT modifications
from a chain truncated by checkchain() prior to reconnection are
written back to disk, requiring "Update FATs? yes" and ensuring a
clean second pass.
Additionally, lost_chain_left_alone, lost_chain_preen, and
corrupted_lost_chain_left_alone cover scenarios that must continue to
report unrecovered errors: read-only mode (-n), which performs no
repairs and leaves the image byte-for-byte unchanged, and preen mode
(-p), which attempts reconnection but does not clear lost chains.
MFC after: 1 week
Invoke releasefat(fat) in checkfilesys() prior to free(fat) on exit paths so that fatbuf, headbitmap.map, and fat32_cache entries are properly freed. MFC after: 1 week Pull Request: https://github.com/freebsd/freebsd-src/pull/2351
derive_mac keeps a per-parent branch index in a global named from the parent interface so the N nibble can increment when the same PHY is presented more than once. That name must be a POSIX identifier; a vlan-style parent (em0.20) is not. Encode the ifname first (alnum unchanged, every other byte as _HH) so the lookup stays a symbol-table hit and em0.20 does not collide with em0_20. Same change in jib (9.2) and jng (9.4). In jng, also address netgraph by node name. ngctl(8) treats `.' and `:' as control characters, so ng_ether(4) names its node after the sanitized ifname (vtnet0.20 becomes vtnet0_20). Sanitize the parent ifname where it enters and use that for every ngctl call; ifconfig(8) and derive_mac keep the real name. Previously jng failed outright on such parents where jib did not. PR: https://bugs.freebsd.org/bugzilla/show_bug.cgi?id=291143 Reported by: Victor <tschetter.victor@gmail.com> MFC after: 1 week Reviewed by: jlduran Differential Revision: https://reviews.freebsd.org/D59326
readboot() decoded the 32-bit little-endian BIOS Parameter Block and FSInfo fields by shifting the individual bytes of a u_char array into place. The u_char operands are promoted to signed int, so shifting a most significant byte of 0x80 or greater left by 24 overflows int, which is undefined behavior. Use le32dec() from <sys/endian.h> instead, which is both well defined and easier to read. No functional change intended. MFC after: 1 week Pull Request: https://github.com/freebsd/freebsd-src/pull/2350
reconnect() computed the byte offset of the LOST.DIR cluster in 32-bit
arithmetic and widened the result only on assignment:
lfoff = (lfcl - CLUST_FIRST) * boot->ClusterSize
+ boot->FirstCluster * boot->bpbBytesPerSec;
cl_t is u_int32_t and ClusterSize is u_int, so both products wrap modulo
2**32. Once LOST.DIR's cluster lies past the 4 GiB mark, lfoff aliases
the offset exactly 4 GiB below it, which on such a volume is ordinary
file data.
That offset is used for both the read and the write: reconnect() reads a
cluster of file data, scans it in 32-byte steps for a leading SLOT_EMPTY
or SLOT_DELETED byte, which arbitrary data readily provides, stores the
new directory entry in that slot, and writes the cluster back to the
same wrong place. Thirty-two bytes of an unrelated file are silently
replaced by a directory entry, and since that entry never reaches the
real LOST.DIR the chain stays lost, so the next run damages another
slot.
Cast to off_t before multiplying. This was the only cluster-to-offset
conversion multiplying a cluster number by the cluster size; the others
in dir.c and fat.c compute a 32-bit sector number first and widen that,
which cannot overflow because the sector count is itself 32-bit.
The bug was observed in the field on a FAT32 stick where LOST.DIR had
been created after a multi-gigabyte file was copied onto it, corrupting
that file every time the volume was checked.
MFC after: 1 week
Pull Request: https://github.com/freebsd/freebsd-src/pull/2347
A HAST message can be empty, in which case ebuf_add_tail() does nothing and ebuf_data() returns NULL because the size of the ebuf is zero, but hast_proto_recv_hdr() asserts that the return value is not NULL, resulting in an immediate crash if hastctl or hastd receive an empty message. This is trivially reproducable by running `hastctl status` or `hastctl role init` (as the rc script does prior to stopping hastd). To avoid this, don't try to grow the ebuf or receive additional data if the header size is zero. PR: https://bugs.freebsd.org/bugzilla/show_bug.cgi?id=298085 MFC after: 3 days Reviewed by: kevans, gjb Differential Revision: https://reviews.freebsd.org/D59306
hastd: Clean up the ebuf code Rename the members of struct ebuf to match their function, replace bcopy() with memcpy(), add comments explaining what each function does. Reviewed by: kevans, emaste Differential Revision: https://reviews.freebsd.org/D59310
hastd: Clean up the ebuf code Missed an instance of bcopy(). Fixes: https://cgit.freebsd.org/src/commit/?id=48c0fc0171a1 ("hastd: Clean up the ebuf code")
gzoned(8) is a new GEOM class that exposes a host-managed zoned device (similar to ZAC/ZBC drives) on top of regular, non-zoned providers. The created medium is sliced into equally sized zones, by default sequential-write-required. Such zones can be turned into conventional zones if desired. The zoned drive's state and configuration is persistent through metadata at the tail of the backing provider, meaning the zoned device gets recreated at the provider retaste. Zone state changes only mark the table dirty with BIO_FLUSH committing it, mirroring drives whose zone state is volatile until a cache flush. The new class tries to emulate real zoned drives by incorporating per-zone write pointers and support for BIO_ZONE management commands. Fault emulation through zone conditions (RWP recommended, offline, R/O), URSWRZ bit toggling and concurrent open zone limits are additional features useful for testing. While at it, fix zonectl's report of write pointer LBAs for zones that should have none, remove its forced debug flags and allow the geom(8) shared subroutines to detect host-managed drives. ATF-sh tests for gzoned(8) and zonectl(8) (dogfooded by gzoned(8)) are included, featuring helpers (see zoned_subr.sh) for testing zoned storage support in other components or GEOM classes. Reviewed by: asomers, fuz, ken Sponsored by: Google Summer of Code 2026 Pull Request: https://github.com/freebsd/freebsd-src/pull/2326
pw_checkfd() returned the character "-" (45) for the "-" argument, which was ambiguous with a real file descriptor. MFC After: 1 week
MFC After: 1 week
pw: cleanup No functional change intended. MFC After: 1 week
pw: fix at job removal when deleting a user rmat() used stat() with a path relative to the current working directory, so it never found the job files in /var/at/jobs and the at(1) jobs of a deleted user were never removed. ef7d0eb9489f also broke it by introducing a typo: /usr/sbin/atrm instead of /usr/bin/artm. Use fstatat() with the directory fd to stat the job files relative to the at jobs directory, and unlinkat() them directly instead of spawning atrm. Those changes allow us to make it works with pw -R. MFC After: 1 week
pw: remove crontab with unlinkat instead of spawning crontab crontab -r only unlinks the crontab file, so spawn it directly with unlinkat() relative to conf.rootfd. This also makes the crontab removal work with pw -R. MFC After: 1 week
pw: remove mail file with unlinkat instead of building a path This is consistent with how at jobs and crontabs are removed. MFC After: 1 week
pw: fix error message in grp_set_passwd to use correct fd MFC After: 1 week
If an explicit loader wasn't requested, then bhyveload(8) maintains a /boot handle that it can use for swapping to a different flavor. This means that we expose all of the host /boot to the sandbox for the duration of script execution. Add a callback to ack that we're OK with the interpreter so that bhyveload(8) can release the bootfd. This is worth doing because it's prior to guest script execution, so we're still running a reasonably untainted process. Reviewed by: imp, jhb Differential Revision: https://reviews.freebsd.org/D58771
MFC After: 1 week
hastd: Ensure nvpair padding is initialized The proto-libnv implementation embedded in hastd pads names and values out to the nearest multiple of eight bytes, but leaves the padding uninitialized, leaking up to 14 bytes of recycled heap per pair in a message. While here, switch from bcopy() to memcpy(). MFC after: 3 days Reviewed by: kevans, emaste Differential Revision: https://reviews.freebsd.org/D59343
hastd: Add a stop control message Add a stop control message which causes hastd to clean up and terminate. Reviewed by: kevans Differential Revision: https://reviews.freebsd.org/D59344
hastd: Add rudimentary tests Test that we can start and stop hastd with an empty configuration. Reviewed by: kevans Differential Revision: https://reviews.freebsd.org/D59345
Discussed with: jrtc27 Sponsored by: The FreeBSD Foundation MFC after: 1 week Differential revision: https://reviews.freebsd.org/D59269
Discussed with: jrtc27 Sponsored by: The FreeBSD Foundation MFC after: 1 week Differential revision: https://reviews.freebsd.org/D59269
Discussed with: jrtc27 Sponsored by: The FreeBSD Foundation MFC after: 1 week Differential revision: https://reviews.freebsd.org/D59269
Discussed with: jrtc27 Sponsored by: The FreeBSD Foundation MFC after: 1 week Differential revision: https://reviews.freebsd.org/D59269
Discussed with: jrtc27 Sponsored by: The FreeBSD Foundation MFC after: 1 week Differential revision: https://reviews.freebsd.org/D59269
Also print the vendor/device and subvendor/subdevice IDs in addition to any strings from the database found if the -v flag is given more than once. This helps with device identification if the strings resolve to identical values for entire product families as well as when the exact card cannot be determined from the string. In theory a second call to pciconf could present that information in non-tree mode but that kind-of defeats the purpose. Reviewed by: jhb MFC after: 3 days Differential Revision: https://reviews.freebsd.org/D56248
Remove struct pmchdr_cpuidinfo which was just a wrapper around a flexible array member of uint32_t. Flexible array members are non-standard in C++, and even in C are not allowed as the only member of a struct. GCC errored out on pmchdr_cpuidinfo, but did not complain about pmchdr_pmcinfo, so I left it alone here, though it is also non-standard. Fixes: https://cgit.freebsd.org/src/commit/?id=93da997ef759 ("pmc: new pmc log processing framework") Reviewed by: Ali Mashtizadeh <ali@mashtizadeh.com> Differential Revision: https://reviews.freebsd.org/D59355
Add NIC-specific VF status to the existing ifconfig -v output. Fetch the data through libifconfig using a separate native route Netlink query. Group optional identity, initialization, resources, VLAN policy, administrator policy, protocol, traffic-permission, and fault containment fields. Omitted fields remain distinct from false or zero. Refer users to iovctl -L for device-neutral PCI attachment and passthrough state. This is a Netlink-native evolution of the original interface by Eric Joyner. Relnotes: yes Sponsored by: Intel Corporation (initial version) Sponsored by: BBOX.io Differential Revision: https://reviews.freebsd.org/D58778 Co-authored-by: Eric Joyner <erj@FreeBSD.org>
Display top-like NFS server I/O using dtrace(1). Also supports JSON output for time-series. Relnotes: yes Reviewed by: ziaee, bcr, adrian Differential Revision: https://reviews.freebsd.org/D59438
When an NSS backend returns a canonical pw_name or gr_name that differs from the lookup name supplied by the NFSv4 upcall, nfsuserd stores the successful mapping in the kernel cache under the canonical name instead of the requested name. This causes the retry lookup performed by nfsv4_strtouid() or nfsv4_strtogid() to miss the newly inserted cache entry, resulting in the default UID/GID being returned although the NSS lookup itself succeeded. PR: https://bugs.freebsd.org/bugzilla/show_bug.cgi?id=296753 Reviewed by: rmacklem Tested by: Emanuel Helms <emanuel@nfv.cc> MFC after: 2 weeks
This fixes the build with gcc 16 after the changes to -Wunused*[0]. [0] https://gcc.gnu.org/gcc-16/porting_to.html#changes-to-wunused Reviewed by: delphij MFC after: 3 days Sponsored by: The FreeBSD Foundation Differential Revision: https://reviews.freebsd.org/D59543
This fixes the build with gcc 16 after the changes to -Wunused*[0]. [0] https://gcc.gnu.org/gcc-16/porting_to.html#changes-to-wunused Reviewed by: glebius MFC after: 3 days Sponsored by: The FreeBSD Foundation Differential Revision: https://reviews.freebsd.org/D59541
This fixes the build with gcc 16 after the changes to -Wunused*[0]. [0] https://gcc.gnu.org/gcc-16/porting_to.html#changes-to-wunused Reviewed by: ken, imp MFC after: 3 days Sponsored by: The FreeBSD Foundation Differential Revision: https://reviews.freebsd.org/D59540
loader: Allow this to compile w/o MD_IMAGE_SIZE defined Slight tweak to allow md.c to be compiled when MD_IMAGE_SIZE isn't defined. Sponsored by: Netflix Reviewed by: olivier, tsoome Differential Revision: https://reviews.freebsd.org/D59415
loader: Small tweak to how md is configured Only the EFI loader supports the md devices (as the moment). Create LOADER_MD_SUPPORT and only try to compile md when that's true. For the moment, also allow it when only MD_IMAGE_SIZE is defined, but that will be dropped since LOADER_MD_SUPPORT is a superset. Sponsored by: Netflix Reviewed by: tsoome Differential Revision: https://reviews.freebsd.org/D59416
loader: Move ramdisk export to md.c Move the common routine to md.c and use it where we do memdisks today. Sponsored by: Netflix Differential Revision: https://reviews.freebsd.org/D59417
loader: Allow dynmaic creation of MD disks Augment the md device to allow creation of md devices. This is useful when UEFI doesn't support RamDisk and in other cases where we have a handy initrd-like image we want to use. Sponsored by: Netflix Differential Revision: https://reviews.freebsd.org/D59418
loader: Use %u to print unsigned value Fixes: https://cgit.freebsd.org/src/commit/?id=7055aa7a070f Sponsored by: Netflix
When someone has loaded a ram disk, that's a stronger hint about what to boot than the UEFI boot manager. Prefer it and rework a little the #ifdefs to reduce their number. Sponsored by: Netflix Differential Revision: https://reviews.freebsd.org/D59419
Allow md devices like efi allows them. Sponsored by: Netflix Differential Revision: https://reviews.freebsd.org/D59420
Some UEFI firmwares do not contain the RamDisk protocol. When this happens, load the file into a md disk instead. This is less desireable than the RamDisk (since that allows better chain-booting options), but sufficient to boot an image as a md device for FreeBSD. While most severs have the RamDisk protocol enabled in their builds, I have two IoT class boards from different manufacturers that do not. Since it works just fine with this fallback, that's preferable. Sponsored by: Netflix Differential Revision: https://reviews.freebsd.org/D59421
Add parse_uri(), a helper for devices whose devspec can take a "[N]://host[:port][/path]" form in addition to the traditional "N:path" form. It returns EINVAL when devspec doesn't match that pattern, so a caller's own dv_parsedev can fall back to the traditional form. Un-static default_parsedev() so it can serve as that fallback. Sponsored by: Netflix Differential Revision: https://reviews.freebsd.org/D59422
stand/efihttp: Support arbitrary http[N]://host[:port]/path devspecs Until now, the "http" device only worked when firmware itself booted loader.efi via EFI HTTP Boot: efihttp_dev_init() required a URI node on the loader's own boot device path, and every fetch went to that same server. Add a second form, using the parse_uri() helper from libsa/dev.c, that works anywhere: http[N]://host[:port]/path (N is a network interface unit, defaulting to 0). This lets a devspec like this be used as rootdev, in loader.conf, or interactively, independent of how loader.efi itself was booted. Network configuration is shared, loader-wide state (the myip/netmask/ gateip/nameip globals from net.h): if something has already configured them (e.g. tftp/nfs via net_open()), the http device just reuses those values, applying them to EFI_IP4_CONFIG2_PROTOCOL as a static address. If nothing has, it drives IP4Config2's own DHCP directly instead of calling netdev's net_open(): net_open() opens the NIC's SNP EFI_OPEN_PROTOCOL_EXCLUSIVE, which would disconnect the Mnp/Ip4/.../ HttpDxe driver chain this code depends on. That chain is also commonly disconnected already by whatever chained us here (e.g. iPXE excludes and exclusively opens SNP for its own raw I/O), so reconnect it explicitly with ConnectController() before giving up. The legacy EFI HTTP Boot path is untouched. Sponsored by: Netflix Differential Revision: https://reviews.freebsd.org/D59423
loader.efi: Fix memory leak in efihttp_dev_close Save enough context to free the host we allocated in open on close. Fixes: https://cgit.freebsd.org/src/commit/?id=6788e42d53c6 Noticed by: claude + Sonet 5 Sponsored by: Netflix
Request FreeBSD RFC 3925 PEM 2238 DHCP 125 option. Use suboption 1 to specify a ramdisk image to load and use to load the kernel and as the root file system. Differential Revision: https://reviews.freebsd.org/D59428 Sponsored by: Netflix
Allow callers to obtain DHCP configuration without opening a network file, enabling loader.efi to discover an initmd URL before handing the interface back to firmware HTTP. Differential Revision: https://reviews.freebsd.org/D59429 Sponsored by: Netflix
loader.efi: release SNP after network configuration Close the exclusively opened SNP protocol after loader networking finishes so the firmware network stack can reconnect the controller and service a subsequent HTTP request. Sponsored by: Netflix Differential Revision: https://reviews.freebsd.org/D59430
loader.efi: download md images through file API Feed the existing decompressor from loader file descriptors as well as iPXE callbacks, allowing arbitrary HTTP filesystem URLs to become kernel-exportable memory disks. Sponsored by: Netflix Differential Revision: https://reviews.freebsd.org/D59431
loader.efi: fetch initmd specified by DHCP Configure net0, read the DHCP-provided initmd URL, and download it through the file API so directly booted loader.efi can construct its root memory disk. Sponsored by: Netflix Differential Revision: https://reviews.freebsd.org/D59432
loader.efi: download DHCP initmd specified image
When we haven't downloaded an image via mem{disk,cd} and ipxe, try any
image specified by the FreeBSD-specific initmd dhcp option. This allows
direct booting of loader.efi with a detached root ramdisk filesystem.
Sponsored by: Netflix
Differential Revision: https://reviews.freebsd.org/D59433
When we can't find root, we give the user a chance to drop to the boot loader OK prompt. We should eat the character that's pressed to not interfere with the interactive session. Sponsored by: Netflix
Hitting <ESC> will now allow one to abort a (possibly mistaken) download. Sponsored by: Netflix
mpsutil: Add Show Discovery command Prints the current discovry tables for the driver. Sponsored by: Netflix Assisted by: Claude (sonet-5) Differential Revision: https://reviews.freebsd.org/D58528
mpsutil: Better naming form the discovery_status function It's really discovery_status_str(). Rename it and use open_memstream() to write the string so we don't have to play as many str* games. Fixes: https://cgit.freebsd.org/src/commit/?id=afb60897a15c Noticed by: claude + Sonet 5 (size error, bad fix ignored) Sponsored by: Netflix
nvmecontrol: Add APST feature support Add -a, -d and -m options to the power subcommand to manage APST, along with additional status output when invoked without args. Signed-off-by: Alexey Sukhoguzov <sap@eseipi.net> Reviewed-by: imp
nvmecontrol: Fix APST transition encoding Use uin64_t casts to encode data sent to avoid overflows. Fixes: https://cgit.freebsd.org/src/commit/?id=35793364d722 Sponsored by: Netflix
nvmecontrol: Minor correctness issues Turn an assert into a bounds check to not overflow if the nvme drive reports too many power states (we validate the user input, but not the drive's identify data). Use a uint32_t instead of int for entry so right shift we do is defined. No functional changes. Fixes: https://cgit.freebsd.org/src/commit/?id=35793364d722 Noticed by: claude + Sonet 5 Sponsored by: Netflix
dhclient(8): Add support for IPv6-Only option (RFC 8925) Accept and validate the ipv6only option. When dhclient receives this option and IPv6 connectivity is available, stop the DHCP configuration process and wait for the duration specified by the option before restarting DHCP discovery. If the address was previously leased, disassociate it and send a DHCPRELEASE packet. Use netlink to check for IPv6 connectivity. Also, unregister ignored options from default PRL. Reviewed by: ziaee, kfv Tested by: Marek Zarychta <zarychtam@plan-b.pwste.edu.pl> Relnotes: yes Differential Revision: https://reviews.freebsd.org/D56637
dhclient(8): Fix dhcpack comment Restore a comment accidentally removed above dhcpack. Fixes: https://cgit.freebsd.org/src/commit/?id=37fd3fcaaf0e ("dhclient(8): Add support for IPv6-Only option (RFC 8925)")
loader.efi: Expose efi_devpath_get_mac to get mac This is a convenient way to test if a device path is a nic or not, so expose it to the world. Sponsored by: Netflix
loader.efi: Only try to download md if we're netbooting The only possible time we could download the initmd that the dhcp server told us about is if we're netbooting. So only attempt to do that if the load device for loader.efi is a network. IF you are booting off disk and then need to snag an initmd off the network, that's a different path, and wouldn't need to necessarily do a dhcp exchange, except to get the IP address. Sponsored by: Netflix
Sponsored by: Netflix
GPT+ZFS+UEFI boots fine, but the manual guided wizard prevented it. MFC after: 3 days Reviewed by: imp, adrian Differential Revision: https://reviews.freebsd.org/D59603
Sponsored by: The FreeBSD Foundation MFC after: 3 days
Incorrect ELF might have PT_NOTE slightly larger than the needed to contain all notes, and the PT_NOTE size could be larger than one page. Then rtld mmaps just the notes bytes to parse. After the last note, we iterate past the mapped region trying to read the Elf_Note header. This was found in wild. Require full elf note to fit into the [start_note, end_note) region to continue the parsing. Check it in stages, first verifying the Elf_Note header structure fits, to be able to read the name and data length. After that, check the whole note against limit. Reported and tested by: makc Reviewed by: emaste Sponsored by: The FreeBSD Foundation MFC after: 1 week Differential revision: https://reviews.freebsd.org/D59635
The bounds check rejected any access ending exactly at the end of the register block, so a four byte write at offset 0xffc was refused, returning EINVAL and killing the VM. MFC after: 1 week PR: https://bugs.freebsd.org/bugzilla/show_bug.cgi?id=291063 Fixes: https://cgit.freebsd.org/src/commit/?id=75909086a45d ("bhyve: allow read/write to full CRB buffer") Sponsored by: Defenso Signed-off-by: Quentin Thébault <quentin.thebault@defenso.fr> Reviewed-by: aokblast, kevans, markj Pull-Request: https://github.com/freebsd/freebsd-src/pull/2362
Reviewed by: fuz, dteske Approved by: fuz (mentor), dteske (mentor) MFC after: 1 week Differential Revision: https://reviews.freebsd.org/D59667
Found with: Claude Code Sonnet 5 MFC after: 2 weeks
mfi_drive_name() built the "Exx:Syy" drive location string using struct mfi_pd_info's encl_index field, the enclosure's firmware- internal position index. Broadcom's own storcli/MegaCli tooling instead leads with the enclosure's Device ID (EID) in its primary drive listing; encl_index only shows up as "Position" in a detailed per-enclosure view. Both numbers are raw, unmodified firmware values already fetched into struct mfi_pd_info/mfi_pd_address, but only encl_index was ever displayed or accepted as input, leading to confusion when cross-referencing drive locations against storcli output. Switch mfi_drive_name() and mfi_lookup_drive() to use encl_device_id instead, aligning FreeBSD's enclosure numbering with Broadcom's own utilities. Since mrsasutil(8) is the same binary as mfiutil(8) under a different name, this applies to both mfi(4) and mrsas(4) alike. This is a user-visible behavior change: the numeric value of "xx" in "Exx:Syy" now differs from before for any enclosure whose EID and position index don't match, affecting anyone scripting against the previous numbering. PR: https://bugs.freebsd.org/bugzilla/show_bug.cgi?id=294353 Reviewed by: imp Relnotes: yes Differential Revision: https://reviews.freebsd.org/D59654
Avoid printing timer addresses, so as to not divulge information about the address space layout. In ip.c, print the actual SPI instead of a pointer to the SPI in the header buffer. When debug logging is enabled, don't leak pointers when logging function arguments or return values. Reported by: Reo Shiseki MFC after: 2 weeks Sponsored by: The FreeBSD Foundation
All communication between hastd nodes and internally between hastd and its worker children passes through the same pair of send / receive functions. The receive function uses recv(2) with the MSG_WAITALL flag, which in theory means we should never get a short read. However, when handing off a socket to a worker child, we also pass a variable-length string identifying the type of socket we're passing, and reading this string relies on a short read. This used to work because the arrival of the descriptor would interrupt the recv(2) call, but this bug was fixed when the AF_UNIX code was rewritten a while ago and hastd has been broken ever since. Fixing the length of the protocol name to four characters including the terminating null solves the short-read bug by never requiring a short read (nothing else in hastd requires one). Note that this issue appears to have been reported independently first by Alessandro Sagratini in PR 234576 and then by Martin Vidovic in D57511. I ended up going in a different direction than Martin's patch, but his analysis was invaluable, hence the double credit below. PR: https://bugs.freebsd.org/bugzilla/show_bug.cgi?id=234576 Reported by: Martin Vidovic <xtronom@gmail.com> MFC after: 1 week Event: EuroBSDcon DevSummit 2026 Reviewed by: xtronom_gmail.com, kevans, glebius, gjb Differential Revision: https://reviews.freebsd.org/D59521
Complete the native configuration trinity: sysctl(8) for live kernel state, sysrc(8) for rc.conf(5), and sysconf(8) for the remaining base configuration -- loader.conf(5), sysctl.conf(5), and the make.conf(5) family -- atop libbsdconf(3). libbsdconf resurrects figpar as a unified reader/writer. Callbacks own semantics; statements may span multiple lines via backslash continuation; non-seekable input is spooled; writes are atomic (mkstemp, fsync, rename) with mode/owner preservation. Format descriptors name each target, its files, and quoting rules without private parsers. Multi-file targets follow boot sourcing order; loader chases loader_conf_files as the boot loader does. sysconf(8) is the operator-facing tool: name / name=value on a required target, sysrc-style list edits, make append and list-strike where they belong, jail/altroot, and a capsicum sandbox for read-only use. Sysctl writes validate against the running kernel first -- unknown and read-only OIDs, CTLFLAG_TUN (pointing at the loader target), and CTLTYPE range checks -- so a typo or overflow does not land in sysctl.conf. Make and src treat WITH_/WITHOUT_ as presence knobs (as bsd.mkopt.mk / src.conf(5) do) and warn on the WITH_*=no form that does not disable the option, so a bad assignment is caught before an /usr/src build surfaces it. The rc target passes through to sysrc(8). Defaults querying (-d/-D/-A) mirrors sysrc for dumps and descriptions on targets that have a defaults file; named reads already see defaults, and -A only widens dump scope. Manuals are split pkg(8)-style (bsdconf/put/format; sysconf plus per-target pages). ATF coverage exercises the frontend. Co-authored-by: Faraz Vahedi <kfv@FreeBSD.org> Reviewed by: fuz, kfv Differential Revision: https://reviews.freebsd.org/D58066
- zero out a buffer for the incoming icmp message that we are about to parse - for ICMP_MASKREPLY and ICMP_TSTAMPREPLY responses check their length and warn and reject them if they are truncated Without these checks a part of stack allocated struct icmp icp could be printed out which seems low severity since we already dropped root privileges by the time icp is allocated. Reviewed by: glebius MFC after: 1 month Found with: Claude Code Sonnet 5 Differential Revision: https://reviews.freebsd.org/D59555
When we see a difference between the payload we sent and what we received, we dump both but we were not prepared for the case when the received payload is less than we sent. In this case we were trying to dump more than needed. Funny enough, 23 years ago I already fixed a similar issue here but didn't pay attention to this small dumping loop. Test written by jlduran. Reviewed by: jlduran MFC after: 1 month Found with: Claude Code Sonnet 5 Differential Revision: https://reviews.freebsd.org/D59556 Differential Revision: https://reviews.freebsd.org/D59582
Change added in 90a7728cd8905cd26b90d06f7873df8bad43ae9a contained a lot of bugs: - If all network interfaces correctly respond to DHCP, then pwait was waiting forever. - If there was more than 1 interface, then "cat /tmp/ephemeraldhcp.*.pid" joined pid numbers into one long string. - "for iface in $left; do kill -15 $left; done" is not using $iface. Tested by: adam.mizerski@ovhcloud.com Sponsored by: OVHcloud Pull Request: https://github.com/freebsd/freebsd-src/pull/2436
Monitor mode will currently unconditionally restart the guest if it reboots, which can be incovenient in certain situations (e.g., resetting the guest after a panic). Address this by introducing a way to prevent automatic guest restarts in monitor mode. Reviewed by: markj Differential Revision: https://reviews.freebsd.org/D59123
Currently, when bhyveload(8) fails to boot the guest, it drops into the loader prompt waiting for user input. This behaviour is inconvenient when using bhyveload(8) from scripts. Make it exit when it receives EOF from console input. PR: https://bugs.freebsd.org/bugzilla/show_bug.cgi?id=286289 Reviewed by: markj MFC after: 2 weeks Sponsored by: The FreeBSD Foundation Differential Revision: https://reviews.freebsd.org/D59226
Reviewed by: fuz Approved by: fuz (mentor) MFC after: 1 week Differential Revision: https://reviews.freebsd.org/D59777
Extend bhyvectl(8) to support querying VM pid using the `--get-vm-pid` flag. This is useful in monitor mode when the VM pid differs from the main bhyve(8) process run by the user. Knowing the VM pid is necessary, for example, to query process resource usage or trigger ACPI shutdown with SIGTERM. Of course, it could be obtained by matching the monitor process children by the process title, but it's a little more complex and fragile than it could be. Implement that by adding the "get_vm_pid" IPC command, and using it to implement `bhyvectl --get-vm-pid`, which prints the VM PID. When the VM PID is not known, ESRCH is returned. Reviewed by: bnovkov (previous iteration without man changes) Sponsored by: The FreeBSD Foundation Differential Revision: https://reviews.freebsd.org/D59108
Add support for dumping the Microsoft Data Management (MSDM) table, which contains the Windows product key embedded in firmware by the OEM. Reviewed by: kib Differential Revision: https://reviews.freebsd.org/D59662
Currently, bhyveload(8) does not validate the supplied disk image path. For example, it allows passing the /dev/null device, which later fails in userboot because it does not support DIOCGSECTORSIZE and DIOCGMEDIASIZE ioctls (see userdisk_init() in stand/userboot/userboot/userboot_disk.c). Fix that by checking DIOCGSECTORSIZE and DIOCGMEDIASIZE ioctls early. A similar check already exists in bhyve(8). While here, make cb_diskioctl() report the obtained sector size instead of hard-coding 512. Reviewed by: markj MFC after: 2 weeks Sponsored by: The FreeBSD Foundation Differential Revision: https://reviews.freebsd.org/D59253
EDK2's QemuBootOrderLib inspects the bootorder file provided via fw_cfg and requires it to be NUL-terminated. Otherwise, it rejects the supplied bootorder and falls back to its default boot order. Currently, bhyve registers bootorder with qemu_fwcfg_add_file() using bootorder_len returned by open_memstream(), which excludes the trailing NUL byte. Fix that by passing bootorder_len + 1 to qemu_fwcfg_add_file() so the fw_cfg payload is properly NUL-terminated. PR: https://bugs.freebsd.org/bugzilla/show_bug.cgi?id=279720 Reviewed by: markj Found with: codex (gpt-5.6-sol) MFC after: 1 week Sponsored by: The FreeBSD Foundation Differential Revision: https://reviews.freebsd.org/D59859
Differential Revision: https://reviews.freebsd.org/D59602 MFC after: 3 weeks PR: https://bugs.freebsd.org/bugzilla/show_bug.cgi?id=289923 Reviewed by: bnovkov
Otherwise, the following would segfault dnctl pipe 1 config bw 1Mbit/s profile 1025points.txt Found with: Claude Code Sonnet 5 MFC after: 2 weeks
Reviewed by: fuz, dteske Approved by: fuz (mentor), dteske (mentor) MFC after: 1 week Differential Revision: https://reviews.freebsd.org/D59876
Check if the response's error_index is within a sane interval. Otherwise, a rogue peer could crash us. PR: https://bugs.freebsd.org/bugzilla/show_bug.cgi?id=298222 Reported by: Robert Morris Reviewed by: markj Discussed with: secteam (markj) MFC after: 2 weeks Analyzed with: Claude Code Opus 5
Otherwise, "bsnmpget -o verbose -I cut" crashes. Reported by: clang static analyzer MFC after: 2 weeks
For the images we read into the memory, use EfiLoaderData memory type instead of EfiReservedMemoryType. The kernel handles the former properly and won't reallocate it. While the latter isn't necessarily mapped, which can cause ram disks to fault when loaded. memdisk_uefi.efi used the latter type, but also registered the device as an ACPI RAM disk, which we don't do. Fixes: https://cgit.freebsd.org/src/commit/?id=afee781523e45 Sponsored by: Netflix
When we're searching for the EFI_GRAPHICS_OUTPUT_PROTOCOL_GUID (GOPs) to use, skip any whose Mode or Mode->Info pointers are NULL. The spec requires these to be non-null, however, some firmwares seem to fail to populate the Info when, for example, a monitor is not present. Work around these bugs by skipping any GOPs with bad pointers. PR: https://bugs.freebsd.org/bugzilla/show_bug.cgi?id=288900 Sponsored by: Netflix Differential Revision: https://reviews.freebsd.org/D59830
Merge the nearly identical level 1 and level 2 filename conversion functions and add output buffer bounds checking. Always NUL-terminate the converted filename, allowing the redundant memset() in cd9660_translate_node_common() to be removed. Signed-off-by: Manuel Einfalt <einfalt1@proton.me> Reviewed by: emaste Pull request: https://github.com/freebsd/freebsd-src/pull/2443
Obtained from: NetBSD 2501d4dfda92, 7022656d8ab2
Obtained from: OpenBSD 4f3c7809d4013a2eb3e75fc5e7bd2d8ae1f41843
There is no need at all to load the cuse module to just access the tool help. kldload always checks for permissions first returning -EPERM if the user can't load modules and -EEXIST if the user can, but the module is already loaded. Approved by: christos@ Differential Revision: https://reviews.freebsd.org/D59844 MFC after: 2 weeks
bitmap_ctor() computed the allocation size as
roundup2(bits, LONG_BIT) / (LONG_BIT / 8)
The dividend is a count of bits, so converting it to bytes requires
dividing by 8 (bits per byte), not by LONG_BIT / 8 (bytes per long).
The two divisors happen to coincide on LP64, but on ILP32 platforms
the head bitmap was allocated at twice the required size; for a
FAT32 file system with close to 2^28 clusters, that is 64 MiB instead
of 32 MiB.
The extra half of the allocation was never accessed, so there is no
functional change other than the reduced memory footprint.
MFC after: 3 days
Pull Request: https://github.com/freebsd/freebsd-src/pull/2440
boot-test.sh: Add virtio-rng-pci to the EFI RAM-disk netboot qemu invocation Since the PixieFail security fixes (CVE-2023-45237), EDK II's DxeNetLib -- underneath essentially all of NetworkPkg (Mnp/Arp/Ip4/Dhcp4/Tcp/Http) -- carries a DEPEX on EFI_RNG_PROTOCOL. With no RNG protocol producer available, that DEPEX is never satisfied and the entire NetworkPkg driver stack silently fails to load: no error, no assert, it just isn't there. netboot-efi and netboot-ramdisk never noticed because they only ever touch the raw EFI_SIMPLE_NETWORK_PROTOCOL via our own net.c, which has no such dependency. A test that needs EDK II's own NetworkPkg (e.g. one exercising EFI_HTTP_PROTOCOL) is the first to be affected. RngDxe can satisfy the DEPEX from the RDRAND instruction alone on a sufficiently recent edk2 build, but not every installed OVMF is that recent. -device virtio-rng-pci provides an RNG unconditionally via VirtioRngDxe, regardless of edk2 vintage or host CPU features. Unfortunately, the edk2 shipped with qemu lacks the network this needs. Sponsored by: Netflix
boot-test.sh: Add a http server per interface Add the built-in python http server, bound to each of the interaces we create to expand network testing to include http:// in various scenarios. Sponsored by: Netflix
boot-test.sh: Add netboot-http-efi test
Using iPXE, chainboot loader.efi with rootdev=http://${next-server}/
to test the loader's http:// code.
Sponsored by: Netflix
BSDRP images are built with "poudriere image -t firmware", which puts
gptboot.efi on the ESP instead of loader.efi. gptboot.efi reads the GPT
bootme attribute, picks the active system partition (BSDRP1 or BSDRP2),
chainloads /boot/loader.efi from it and hands loader.efi that partition
in LoadedImage->DeviceHandle (stand/efi/boot1/boot1.c:try_boot).
Since ce9bfd78167 ("loader.efi: Refactor try_boot_device_partitions"),
find_currdev() no longer tries that device: try_boot_device_partitions()
walks the parent disk while explicitly skipping dp->pd_handle, on the
assumption that the boot image always comes from an ESP holding no root
filesystem. When chainloaded, the partition gptboot.efi selected is
therefore the one partition never considered, and loader.efi falls
through to the first other UFS partition on the disk - the previous
system. Every A/B upgrade silently boots the old slice.
Commit 1c85c5eea09, which introduced try_boot_device_partitions(), did
try dp itself before its siblings; the refactor dropped it. Restore it,
keeping the sibling walk as the fallback for the normal ESP case.
Fixes: https://cgit.freebsd.org/src/commit/?id=ce9bfd78167
Assisted-by: Claude Code (Fable 5, Opus 5)
Sponsored by: Netflix
Make sure that we can chainboot with gptboot.efi. Many projects use this as their migration tool from gptboot to ping-pong partitions to boot from. This tests that functionality which I recently broke. Assisted-by: Claude Code (Fable 5, Opus 5) Sponsored by: Netflix
Report net.ue.<unit>.%parent as "parent interface", consistent with
wlan(4):
# ifconfig ue0
ue0: flags=8843<UP,BROADCAST,RUNNING,SIMPLEX,MULTICAST> metric 0 mtu 1500
ether 00:00:5e:00:53:2a
inet 192.0.2.10/24 broadcast 192.0.2.255
parent interface: ure0
media: Ethernet autoselect (2500Base-T <full-duplex>)
status: active
Reviewed by: aokblast
MFC after: 2 weeks
Sponsored by: The FreeBSD Foundation
Differential Revision: https://reviews.freebsd.org/D59060
When the OpenFirmware loader flattens the firmware device tree into the FDT it hands to the kernel (usefdt=1, i.e. on every real-mode OF system such as pSeries LPARs and QEMU pseries guests), add_node_to_fdt() clamps every property value to 1024 bytes. Any larger property reaches the kernel truncated. On QEMU pseries the PCI host bridge's "interrupt-map" is 3584 bytes (32 slots x 4 pins x 7 cells), so only the entries for slots 0-8 survive and the entry for slot 9 is cut in the middle. A PCI device in slot 9 or above therefore gets no INTx routing (irq 0), and with INVARIANTS the partial trailing entry trips the "ofw_bus_search_intrmap: truncated map" assertion in ofw_bus_search_intrmap() during PCI attach, panicking the kernel as soon as such a device is present. "ibm,drc-indexes", "ibm,drc-names" and "ibm,drc-power-domains" are cut the same way. Drop the clamp. fdt_setprop() already reports a property that does not fit into the FDT buffer, so no separate limit is needed. Fixes: https://cgit.freebsd.org/src/commit/?id=5ecf8e3852a9 ("Clean up some FDT-related code in the PowerPC bootloader, improving error checking and robustness. Prevents errors and crashes in FDT commands on PowerMac G5 systems.") Differential Revision: https://reviews.freebsd.org/D60086
Ability to override DHCP options with command-line arguments was affected by intoduction of initmd support. Initmd discovery configures the network before loader arguments were parsed, so a dhcp.root-path override was unavailable during the first network configuration. Move parsing arguments earlier and apply dhcp.root-path even if DHCP response does not contain option 17. This allows providing a dynamic NFS root e.g. by chain loading loader.efi from iPXE. Signed-off-by: Krzysztof Galazka <krzysztof.galazka@intel.com> Reviewed by: imp Assisted by: Github Copilot (GPT-5.6 Sol) Sponsored by: Intel Corporation Differential Revision: https://reviews.freebsd.org/D59998
DHCP initmd discovery used net_configure() to open and configure SNP, then immediately closed the socket and released the protocol. Selecting net0 as currdev subsequently required another exclusive SNP attach, which may hang on some EDK II firmware. Keep the configured socket available for ordinary network consumers so NFS boot can reuse it. Add an explicit deconfiguration operation and invoke it only when an initmd URL requires handing SNP back to the firmware network stack. Signed-off-by: Krzysztof Galazka <krzysztof.galazka@intel.com> Reviewed by: imp Assisted by: Github Copilot (GPT-5.6 Sol) Sponsored by: Intel Corporation Differential Revision: https://reviews.freebsd.org/D59999
pci_vtcon_announce_port() sends VIRTIO_CONSOLE_PORT_NAME immediately after VIRTIO_CONSOLE_DEVICE_ADD, without waiting for the guest to initialize the port. The Windows virtio-serial driver ignores the name if the port is not yet available, which can prevent the QEMU guest agent from finding org.qemu.guest_agent.0. Send the name after the guest reports successful port initialization with VIRTIO_CONSOLE_PORT_READY, as QEMU does. While here, validate the guest-supplied port ID before forming a pointer into vsc_ports[]. This ports the PORT_NAME ordering fix from illumos change 18082. Obtained from: illumos 643ee887be44404b690eeaff14f20a27ad3c715c Reviewed by: markj MFC after: 2 weeks Differential Revision: https://reviews.freebsd.org/D59847
paddr_guest2host() returns NULL when a range does not fit within guest RAM. An unmappable queue address previously left the queue marked allocated with ring pointers computed from the failed mapping. Check the legacy virtqueue ring mapping before updating the queue state, and ignore an unmappable PFN write after printing a diagnostic. This leaves an unallocated queue unallocated and preserves an existing queue's PFN, ring pointers, flags, and indices. Also check the indirect descriptor table mapping before dereferencing it, returning an error if the table is unmappable, and ignore guest queue notifications for unallocated queues. Reviewed by: markj MFC after: 2 weeks Sponsored by: The FreeBSD Foundation Differential Revision: https://reviews.freebsd.org/D59904
c2d03a920ec7 rewrote the action printing in print_rule() from OpenBSD, which has no natpass, and dropped the "pass" keyword. A ruleset loaded from "pfctl -sn" output therefore lost its nat-pass semantics. Add a parser test covering nat, rdr, rdr log and binat with pass. Reviewed by: kp Approved by: kp (mentor) Fixes: https://cgit.freebsd.org/src/commit/?id=c2d03a920ec7 ("pfctl: fix anchortypes bounds test") MFC after: 1 week Sponsored by: Rubicon Communications, LLC ("Netgate") Differential Revision: https://reviews.freebsd.org/D60183
Fixes: https://cgit.freebsd.org/src/commit/?id=255538cd906045095d0c2113ae6c4731ce36c0cf Differential Revision: https://reviews.freebsd.org/D57850 Reviewed by: adrian
MFC after: 1 week Fixes: https://cgit.freebsd.org/src/commit/?id=c3276e02beab ("sockets: make shutdown(2) how argument a enum") Reviewed by: glebius Differential Revision: https://reviews.freebsd.org/D57915
PR: https://bugs.freebsd.org/bugzilla/show_bug.cgi?id=296410 Submitted by: Tomas Vondra <tomas@vondra.me> MFC after: 1 week
libsysdecode: decode Generic Netlink controller messages Decode Generic Netlink controller (GENL_ID_CTRL) messages in Netlink payloads. Display the Generic Netlink header along with the CTRL_CMD_GETFAMILY attributes, including the family ID and family name. Signed-off-by: Ishan Agrawal <iagrawal9990@gmail.com> Reviewed by: kp Sponsored-by: Google LLC (GSoC 2026)
libsysdecode: cache Generic Netlink family IDs Record Generic Netlink family IDs learned from CTRL_CMD_GETFAMILY responses and use them to decode subsequent Generic Netlink messages using symbolic family names instead of numeric IDs. Signed-off-by: Ishan Agrawal <iagrawal9990@gmail.com> Reviewed by: kp Sponsored-by: Google LLC (GSoC 2026)
libsysdecode: decode PF Generic Netlink commands Decode the Generic Netlink command header for messages belonging to the PF Generic Netlink family. Display the command name using the PF Generic Netlink command decoder. Signed-off-by: Ishan Agrawal <iagrawal9990@gmail.com> Reviewed by: kp Sponsored-by: Google LLC (GSoC 2026)
libsysdecode: add attribute parsing for PFNL_CMD_GETRULES Signed-off-by: Ishan Agrawal <iagrawal9990@gmail.com> Sponsored-by: Google LLC (GSoC 2026) Reviewed by: kp
libsysdecode: use decoder table for PF netlink commands Introduce a PF netlink command decoder table mapping PFNL commands to their attribute decoder sets. This replaces the existing switch-based dispatch and makes it easier to add support for additional PF netlink commands. Signed-off-by: Ishan Agrawal <iagrawal9990@gmail.com> Sponsored-by: Google LLC (GSoC 2026) Reviewed by: kp
libsysdecode: avoid extra commas for undecoded netlink attributes Signed-off-by: Ishan Agrawal <iagrawal9990@gmail.com> Sponsored-by: Google LLC (GSoC 2026) Reviewed by: kp
libsysdecode: verify decoder tables are sorted Add assertions to validate decoder table ordering required by binary search. Signed-off-by: Ishan Agrawal <iagrawal9990@gmail.com> Sponsored-by: Google LLC (GSoC 2026) Reviewed by: kp
libc/resolv: Drop Solaris 2 compatibility MFC after: 1 week Reviewed by: kevans, markj Differential Revision: https://reviews.freebsd.org/D57922
libc/resolv: Refactor the option parser Start the loop by finding the end of the option name, the name-value separator (if any), and the end of the option. Use those pointers to simplify matching the option name and parsing the option value, and validate option names and values more strictly. This means that: * We no longer accept trailing garbage in an option name or value. For instance, we would previously interpret “edns0123” as “edns0” and “timeout:3xyz” as “timeout:3”. This was actually quite lucky because we also failed to recognize the newline at the end of the option line as a whitespace character. * For options that take a numerical argument, we would previously accept negative values and treat non-numerical arguments as 0, while large numerical arguments would be capped to the option's maximum permitted value. Now, any failure to parse the argument, including overflow, results in the option being left unchanged. MFC after: 1 week Relnotes: yes Reviewed by: markj Differential Revision: https://reviews.freebsd.org/D57923
libc/resolv: Refactor the configuration parser This was previously all a single loop in res_init(), apart from option parsing which we cleaned up in a previous commit. Break it out into separate functions for reading the configuration line by line, setting the default domain, setting the search list, and adding a nameserver to the nameserver list. Sprinkle bounds checks and code comments all around. The sortlist code, which has been disabled for the past 20 years, will be dealt with in a separate commit. MFC after: 1 week Reviewed by: markj Differential Revision: https://reviews.freebsd.org/D57924
libc/resolv: Reimplement the sortlist parser When we switched from the BIND4 resolver to the BIND9 resolver, the sortlist parser was inadvertently disabled due to a missing #define, and nobody seemed to notice. The sorting code remained enabled in the resolver, but there was no way to set a sort order. Reimplement the sortlist parser, but correctly, and update the manual accordingly. The new parser accepts IPv4 and IPv6 addresses with or without a mask or prefix length, just like the old one, except IPv6 support was a bit wonky in the original code. Fixes: https://cgit.freebsd.org/src/commit/?id=5342d17f09a8 ("Update the resolver in libc to BIND9's one.") Relnotes: yes Reviewed by: markj Differential Revision: https://reviews.freebsd.org/D57925
libc/resolv: Add no-debug and no-rotate options These are simply the reverse of the debug and rotate options. Reviewed by: markj Differential Revision: https://reviews.freebsd.org/D57926
Reviewed by: markj Differential Revision: https://reviews.freebsd.org/D57927
getfacl / acl_to_text() incorrectly prints uid/gid numbers as signed integers. This causes uid / gid numbers larger than 2G (2147483648) to print as negative numbers. The libc acl_from_text() function does not handle negative numbers. This diff adds a backwards compatiblity fix to allow negative numbers... Reviewed by: rmacklem MFC after: 2 weeks Differential Revision: https://reviews.freebsd.org/D57180
While here, clean up and simplify the existing code. MFC after: 1 week Reviewed by: glebius, jhb Differential Revision: https://reviews.freebsd.org/D57993
Reviewed by: markj Tested by: pho Sponsored by: The FreeBSD Foundation MFC after: 1 week Differential revision: https://reviews.freebsd.org/D57124
Reviewed by: markj Tested by: pho Sponsored by: The FreeBSD Foundation MFC after: 1 week Differential revision: https://reviews.freebsd.org/D57163
libfetch: Overhaul socket read / write * Make fetch_ssl_read() and fetch_ssl_write() behave more like read(2) and write(2), and drop fetch_socket_read() in favor of read(2). * Don't request POLLERR, it's implied. * Don't needlessly set errno, it's relatively costly. * Always check for EAGAIN from writev(2), otherwise we will abort on a short write instead of proceeding to poll(2). * Always check for EAGAIN from poll(2) even though it can't happen on FreeBSD; POSIX says it can, and it might in the future. * Rewrite fetch_read() and fetch_writev() to be more similar to each other. The main difference is that a partial read is treated as success while a partial write is treated as failure. PR: https://bugs.freebsd.org/bugzilla/show_bug.cgi?id=296316 MFC after: 1 week Reviewed by: op Differential Revision: https://reviews.freebsd.org/D57906
libfetch: Add read buffering Previously, we would read FTP control connection messages and HTTP reponse headers one character at a time. Now, we read as much as will fit in our buffer and look for a newline. If there is data left over, it will be reused by the next fetch_getln() call. This also requires the addition of a fetch_bufread() which takes the buffer into account, otherwise the start of the HTTP response body will be stuck in the buffer after we read the last line of the header. This should noticeably improve HTTP performance, especially for small transfers. MFC after: 1 week Reviewed by: op Differential Revision: https://reviews.freebsd.org/D57907
libfetch: Apply timeout to connection attempts Mark the socket non-blocking before connecting and poll for completion, applying fetchTimeout if set. MFC after: 1 week Reviewed by: op Differential Revision: https://reviews.freebsd.org/D57909
libfetch: Limit response line length When reading an FTP or HTTP response, error out if we read 64 kB before hitting a newline. Otherwise a runaway or malicious server could have us spinning for quite a while allocating more and more memory before we gave up or crashed. MFC after: 1 week Reviewed by: markj Differential Revision: https://reviews.freebsd.org/D60007
libfetch: Reject control characters in credentials If a URL contains credentials, check that neither the user name nor the password contain control characters which might confuse the server, especially in the FTP case. MFC after: 1 week Reviewed by: markj Differential Revision: https://reviews.freebsd.org/D60008
libfetch: Correctly free redirect target When processing a redirect, if we get multiple Location headers, we free the previous one using free(new) instead of fetchFreeURL(new), which leaks new->doc. While here, rename new to loc since new is a reserved word in C++. MFC after: 1 week Reviewed by: markj Differential Revision: https://reviews.freebsd.org/D60015
libfetch: Plug leak in 304 Not Modified case When we get a 304 Not Modified response, we need to go through the error path to properly free any resources we've allocated instead of just returning NULL directly. MFC after: 1 week Reviewed by: markj Differential Revision: https://reviews.freebsd.org/D60016
libfetch: Plug connection leaks in FTP code When setting up an FTP transfer, we need to dereference the cached connection before returning after a failure. MFC after: 1 week Reviewed by: markj Differential Revision: https://reviews.freebsd.org/D60017
Make fetch_ref() an inline and provide a fetch_deref(). MFC after: 1 week Reviewed by: op Differential Revision: https://reviews.freebsd.org/D57944
Reduce the amount of copying we do when performing buffered reads. MFC after: 1 week Reviewed by: op Differential Revision: https://reviews.freebsd.org/D58113
Reviewed by: zlei, vmaffione Obtained from: https://github.com/luigirizzo/netmap/commit/b52a2bcae35e56548acfb0849b248a1e4b0c0c3b MFC after: 2 weeks Differential Revision: https://reviews.freebsd.org/D58150
Reviewed by: zlei, vmaffione Obtained from: https://github.com/luigirizzo/netmap/commit/7d9177ed9a121e66bf4eaa0acb5d574e408297da MFC after: 2 weeks Differential Revision: https://reviews.freebsd.org/D58151
Logic prior to this change would incorrectly try linking when MK_CDDL != no, instead of MK_CTF != no, which could result in the library and the tests being broken if/when MK_CTF == no and MK_CDDL != no (an uncommon, but possible combination with today's build knobs). This change updates the conditional to correctly track the value of MK_CTF, which in turn is properly toggled to no if/when MK_CDDL == no as it's a dependent build knob. This [niche] build bug has been present in FreeBSD since 2014. MFC after: 1 week
* Get rid of the pointless LOCALBASE_CTL_LEN mechanism * Apply minimal normalization to the paths obtained from the environment or sysctl variable * Turn the manual page into a manual page * Add tests MFC after: 1 week Reviewed by: se Differential Revision: https://reviews.freebsd.org/D58362
MFC after: 1 week Fixes: https://cgit.freebsd.org/src/commit/?id=2a5e58c59694 ("procdesc: add NOTE_PDSIGCHLD") Reviewed by: kib Differential Revision: https://reviews.freebsd.org/D58388
Introduce a generic Netlink attribute decoding framework based on attribute decoder tables. The framework supports decoding primitive attribute types as well as nested attributes and can be reused by different Generic Netlink families. Signed-off-by: Ishan Agrawal <iagrawal9990@gmail.com> Sponsored-by: Google LLC (GSoC 2026) Reviewed-by: kp Pull-Request: https://github.com/freebsd/freebsd-src/pull/2337
Sponsored by: The FreeBSD Foundation
Sponsored by: The FreeBSD Foundation
Replace <sysdecode.h> with "sysdecode.h" so local builds use the in-tree header in lib/libsysdecode instead of a stale installed copy in /usr/include, avoiding build failures after updating sysdecode.h. Signed-off-by: Ishan Agrawal <iagrawal9990@gmail.com> Sponsored-by: Google LLC (GSoC 2026) Reviewed by: kp Pull-Request: https://github.com/freebsd/freebsd-src/pull/2338
Change sysdecode_nlm_flag() to decode Netlink message flags as a bitmask instead of looking up a single flag value. This correctly prints combinations of NLM_F_* flags while preserving any unknown bits in hexadecimal. Reported by: androvonx95 <androvonx95@tutamail.com> Reviewed by: kp Signed-off-by: Ishan Agrawal <iagrawal9990@gmail.com> Sponsored-by: Google LLC (GSoC 2026)
Reviewed by: markj Sponsored by: The FreeBSD Foundation MFC after: 1 week Differential revision: https://reviews.freebsd.org/D58463
libfetch: Fix handling of connection failures After commit 848f360c8f9a, if one tries to connect to a closed port, fetch reports "Operation now in progress", which is rather confusing. Return a more useful error message, restoring the old behaviour. Fixes: https://cgit.freebsd.org/src/commit/?id=848f360c8f9a ("libfetch: Apply timeout to connection attempts") Reviewed by: des MFC after: 3 days Differential Revision: https://reviews.freebsd.org/D58481
libfetch: Further improve connection polling * Reorganize the connection loop to make it a little more readable * Start the timeout clock earlier * Correctly calculate the poll timeout before calling poll() * Don't leak the socket on failure Fixes: https://cgit.freebsd.org/src/commit/?id=848f360c8f9a ("libfetch: Apply timeout to connection attempts") Fixes: https://cgit.freebsd.org/src/commit/?id=b02e02958dad ("libfetch: Fix handling of connection failures") MFC after: 3 days Reviewed by: markj Differential Revision: https://reviews.freebsd.org/D58512
Some NLM_F_ definitions contain multiple underscores in their name; this should pick them up. Reviewed by: kp, Ishan Agrawal <iagrawal9990@gmail.com> Fixes: https://cgit.freebsd.org/src/commit/?id=4c932a4d45fb ("netlink: decode netlink message flags symbolically") Sponsored by: The FreeBSD Foundation Pull Request: https://github.com/freebsd/freebsd-src/pull/2340
Reviewed by: jrtc27 Differential Revision: https://reviews.freebsd.org/D58558
stdio: *memstream: slightly streamline growth function Inverting the condition after realloc*() is a minor cleanup, but makes the success path a little cleaner to ease a future change. Reviewed by: des, jhb Sponsored by: Klara, Inc. Differential Revision: https://reviews.freebsd.org/D57353
stdio: *memstream: decouple the buffer size from the stream length It's useful to be able to track both facts with a single variable, but it also makes it more difficult to change how the buffer size scales. As an example, Apple's implementation seems to scale the buffer size by 1.5x on growth, presumably in an attempt to reduce trips into realloc(). This might be questionable in the face of stdio buffering, but avoiding serious churn in the line- or un-buffered case is a net positive if doing so isn't incredibly invasive. Reviewed by: des, jhb, obiwac Sponsored by: Klara, Inc. Differential Revision: https://reviews.freebsd.org/D57354
stdio: *memstream: grow the buffer by 1.5x on write This improves performance by reducing the number of allocations as we write into the memstream, both in the fully buffered case with larger memstreams and also more trivially in the line- and un-buffered case as they flush back to the underlying buffer more often. The inspiration for this was taken from Apple's implementation in https://github.com/apple-oss-distributions/libc, but expanded to include wmemstream for consistency. I've added a test for the bug that I hit in libder that caused me to notice this in the first place, and fixed that bug in this version. Reviewed by: des, jhb (both slightly previous version) Sponsored by: Klara, Inc. Differential Revision: https://reviews.freebsd.org/D57355
stringf.cc uses errno and related macros without including <cerrno>. Their availability is guaranteed only when the corresponding header is included; transitive exposure is implementation-defined. Modern libc++ has been progressively reducing incidental transitive includes as part of its header removal policy (see LLVM libc++ Header Removal Policy and D132284), making such dependencies brittle. This change includes <cerrno> explicitly to make the dependency well-defined. No functional or behavioural change intended. Approved by: fuz Signed-off-by: Faraz Vahedi <kfv@kfv.io> Pull-Request: https://github.com/freebsd/freebsd-src/pull/2188
Currently mergesort() uses ICOPY_*() to copy data as four byte blocks instead of one byte. However, this is only achievable when both size and base arguments are aligned to four bytes. Use of memcpy() is ideal as 1) it is cleaner and 2) the library will use SIMD for copying when the hardware supports it. Compared to ICOPY_*(), SIMD can support up to 64 bytes. When the SIMD-backed memcpy() find the address is unaligned, it can first copy data up to the nearest aligned address, and then use SIMD operations for faster transfer. Thus memcpy() can give better performance than mergesort()'s own implementation. This is benchmarked on amd64 where there isn't a SIMD-backed implementation yet. However, the baseline implementation in assembly already delivers better performance in unaligned cases although there is some performance drops in aligned cases. The benchmark results and script is available in the Phabricator review. Ideally, more performance improvements will come when amd64 gets SIMD implementation of memcpy(). Signed-off-by: Minsoo Choo <minsoochoo0122@proton.me> Reviewed by: fuz MFC after: 1 week Differential Revision: https://reviews.freebsd.org/D58002
This commit implements the inverse half-cycle
trigonometric functions:
asinpi(x) = asin(x) / pi Eq. (1)
acospi(x) = acos(x) / pi
atanpi(x) = atan(x) / pi
Implemention details are contained in src/s_asinpi.c and
src/a_atanpi.c, where the details for acospi(x) appear in
the former.
*************
CAVEAT EMPTOR: The ld128 code has been only compiled. It has
not been tested for correctness due to lack of hardware.
*************
Code compiled on AMD Ryzen 7 7700X system run FreeBSD 16.0-CURRENT
main-n284956-de9fe28ab847.
Exhaustive testing of acospif(x), asinpif(x), and atanpif(x)
on the indicated intervals yields
% ./tlibm acospi -fPE -x 0x1p-120 -X 1
Interval tested for acospif: [7.52316e-37,1]
ulp <= 0.5: 99.627% 1002878299 | 99.627% 1002878299
0.5 < ulp <= 0.6: 0.277% 2789599 | 99.904% 1005667898
0.6 < ulp <= 0.7: 0.096% 965062 | 100.000% 1006632960
Max ulp: 0.63661975 at 5.96046412e-08 0x1.fffffep-25
% ./tlibm asinpi -fPED -x 0x1p-120f -X 1.f
Interval tested for asinpif: [7.52316e-37,1]
ulp <= 0.5: 99.851% 1005129353 | 99.851% 1005129353
0.5 < ulp <= 0.6: 0.149% 1501097 | 100.000% 1006630450
0.6 < ulp <= 0.7: 0.000% 2510 | 100.000% 1006632960
Max ulp: 0.68957579 at 5.04878759e-01 0x1.027f78p-1
% ./tlibm atanpi -fPE -x 0x1p-120 -X max > zatanpif.txt &
Interval tested for atanpif: [7.52316e-37,3.40282e+38]
ulp <= 0.5: 99.865% 2077574602 | 99.865% 2077574602
0.5 < ulp <= 0.6: 0.131% 2735011 | 99.997% 2080309613
0.6 < ulp <= 0.7: 0.003% 65170 | 100.000% 2080374783
Max ulp: 0.68433094 at 5.01186252e-01 0x1.009b7cp-1
Testing the double and long double version cannot be done
in an exhaustive manner. For 300 M values, uniformily
distributed in the indicated interals, one finds the max ULP:
Interval tested for acospi: [9.31323e-10,0.25]
xm = 2.4423788416892520e-01, /* 0x3fcf432f, 0xde79920f */
libm = 4.2146222480005391e-01, /* 0x3fdaf93c, 0xb201001c */
mpfr = 4.2146222480005396e-01, /* 0x3fdaf93c, 0xb201001d */
ULP = 0.50499351466286857
Interval tested for acospi: [0.25,0.5]
xm = 4.9689430915631438e-01, /* 0x3fdfcd1d, 0xc9d945c6 */
libm = 3.3447366122373884e-01, /* 0x3fd56804, 0x371513ef */
mpfr = 3.3447366122373889e-01, /* 0x3fd56804, 0x371513f0 */
ULP = 0.57195275455053829
Interval tested for acospi: [0.5,0.75]
xm = 5.0238623667462079e-01, /* 0x3fe0138c, 0x4d0f4be0 */
libm = 3.3245556599062825e-01, /* 0x3fd546f3, 0xb5d36303 */
mpfr = 3.3245556599062820e-01, /* 0x3fd546f3, 0xb5d36302 */
ULP = 0.63427929243758807
Interval tested for acospi: [0.75,1]
xm = 7.5853651919512177e-01, /* 0x3fe845ee, 0x60d8789f */
libm = 2.2591472240382732e-01, /* 0x3fcceac6, 0x0c3465ce */
mpfr = 2.2591472240382729e-01, /* 0x3fcceac6, 0x0c3465cd */
ULP = 0.56915750216472161
Interval tested for asinpi: [9.31323e-10,0.25]
xm = 1.9502362835488171e-01, /* 0x3fc8f688, 0xc4dda0fb */
libm = 6.2478354989018887e-02, /* 0x3faffd29, 0xb6c57c61 */
mpfr = 6.2478354989018881e-02, /* 0x3faffd29, 0xb6c57c60 */
ULP = 0.52347765415885006
Interval tested for asinpi: [0.25,0.5]
xm = 4.9937103583123676e-01, /* 0x3fdff5b1, 0xeeddbf62 */
libm = 1.6643553767987129e-01, /* 0x3fc54dc2, 0x7b9d15a4 */
mpfr = 1.6643553767987126e-01, /* 0x3fc54dc2, 0x7b9d15a3 */
ULP = 0.66214688371031072
Interval tested for asinpi: [0.5,0.75]
xm = 5.0228515250761718e-01, /* 0x3fe012b8, 0x4fe92bbb */
libm = 1.6750722213679006e-01, /* 0x3fc570e0, 0x6c75edd5 */
mpfr = 1.6750722213679009e-01, /* 0x3fc570e0, 0x6c75edd6 */
ULP = 0.78223048105528226
Interval tested for asinpi: [0.75,1]
xm = 7.5425933001419776e-01, /* 0x3fe822e4, 0x7663a4aa */
libm = 2.7200385380185182e-01, /* 0x3fd16882, 0xda1dc13b */
mpfr = 2.7200385380185188e-01, /* 0x3fd16882, 0xda1dc13c */
ULP = 0.53747973176773822
Interval tested for atanpi: [9.31323e-10,0.25]
xm = 1.9666113418757322e-01, /* 0x3fc92c31, 0x29dd6d2f */
libm = 6.1810387818117797e-02, /* 0x3fafa59c, 0x7476baa5 */
mpfr = 6.1810387818117804e-02, /* 0x3fafa59c, 0x7476baa6 */
ULP = 0.54674297446584263
Interval tested for atanpi: [0.25,0.5]
xm = 4.1312119637707068e-01, /* 0x3fda7093, 0xe2ee5494 */
libm = 1.2470309560460152e-01, /* 0x3fbfec8a, 0xc554ebec */
mpfr = 1.2470309560460154e-01, /* 0x3fbfec8a, 0xc554ebed */
ULP = 0.73116638175113347
Interval tested for atanpi: [0.5,0.75]
xm = 5.0018949583396499e-01, /* 0x3fe0018d, 0x66cd1b82 */
libm = 1.4763186871058706e-01, /* 0x3fc2e599, 0xdffacb8f */
mpfr = 1.4763186871058709e-01, /* 0x3fc2e599, 0xdffacb90 */
ULP = 0.69192753950764663
Interval tested for atanpi: [0.75,1]
xm = 7.5007880583359599e-01, /* 0x3fe800a5, 0x448f4c03 */
libm = 2.0484881828445453e-01, /* 0x3fca387c, 0x6f93f71f */
mpfr = 2.0484881828445450e-01, /* 0x3fca387c, 0x6f93f71e */
ULP = 0.65765471872064396
Interval tested for atanpi: [1,2]
xm = 1.0103228000344093e+00, /* 0x3ff02a48, 0x3d88d0a2 */
libm = 2.5163447403817019e-01, /* 0x3fd01ac7, 0x7b229108 */
mpfr = 2.5163447403817013e-01, /* 0x3fd01ac7, 0x7b229107 */
ULP = 0.67409519689166042
Interval tested for atanpi: [2,4]
xm = 2.0231383267437946e+00, /* 0x40002f63, 0x25a530a9 */
libm = 3.5387589538123299e-01, /* 0x3fd6a5e7, 0x156053c6 */
mpfr = 3.5387589538123293e-01, /* 0x3fd6a5e7, 0x156053c5 */
ULP = 0.69695587476021503
Interval tested for atanpi: [4,1.79769e+308]
xm = 4.0000000000000000e+00, /* 0x40100000, 0x00000000 */
libm = 4.2202086962263069e-01, /* 0x3fdb0263, 0xd2508e31 */
mpfr = 4.2202086962263069e-01, /* 0x3fdb0263, 0xd2508e31 */
ULP = 0.27709400511686716
PR: https://bugs.freebsd.org/bugzilla/show_bug.cgi?id=295884
MFC after: 1 month
Reviewed by: fuz
strfromd(), strfromf(), and strfroml() are implemented directly in terms of gdtoa. If a non-conforming format string is passed, the string "EDOOFUS" is returned and errno set to EDOOFUS as an extension. Reviewed by: fuz MFC after: 1 month Pull-Request: https://github.com/freebsd/freebsd-src/pull/2301 Signed-off-by: Faraz Vahedi <kfv@kfv.io>
On some platforms, e.g. Linux Clang 22.1.8 / glibc 2.43, strchr() now implements the C23 behaviour where passing a const pointer to strchr() also returns a const pointer. This breaks getopt during the bootstrap build, since it assumes the return value is always a mutable pointer. Since the pointed-to value is never modified, fix this by making the pointer const. MFC after: 1 week Reviewed by: emaste Sponsored by: The FreeBSD Foundation Differential Revision: https://reviews.freebsd.org/D58488
Replace all _open() calls with _openat() in __fts_open(), fts_read(), and fts_children(). Replace statfs() with _fstatfs(). Add fts_dirfd to struct _ftsent, set to the file descriptor of the parent directory. Callers can use openat(ent->fts_dirfd, ent->fts_name, ...) to access files safely without relying on fts_accpath, which enables programs in capability mode to open the files described by _ftsent. This is a preparatory change for fts_openat() which will allow callers to provide a pre-opened directory fd, enabling fts(3) traversal inside Capsicum capability mode. Mirror all fts_open() changes to fts_open_b(). As a result of expanding _ftsend, publish new ELF symbol versions for fts_openat and related functions. Sponsored by: Google LLC (GSoC 2026) Reviewed by: asomers Pull Request: https://github.com/freebsd/freebsd-src/pull/2303
Add an uncached sysctl based implementation which retrieves individual categories. A cache would be an obvious extension should this optional feature that can only be enabled by an environmental varible have a noticable performance impact in a case that matters. Reviewed by: kib Sponsored by: Innovate UK Differential Revision: https://reviews.freebsd.org/D58238
When fts_children() is called, sp->fts_child is set. On the next fts_read() call, fts_safe_changedir() was incorrectly passed p->fts_dirfd (pointing to the parent directory) instead of -1. This caused fts to fchdir to the parent instead of the child directory, silently skipping the contents of 3rd-level subdirectories. This was observed as a failure in nmtree_test:mtree_create which calls fts_children() internally. Add regression test accpath_correct_after_descent that calls fts_children() on each directory entry and verifies files at depth 3 are still visited correctly. The test fails with the unfixed libc and passes with the fix. Reported by: Herbert J. Skuhra <herbert@gojira.at> Sponsored by: Google LLC (GSoC 2026) Reviewed by: asomers Fixes: https://cgit.freebsd.org/src/commit/?id=4bd01d6ae01 ("fts: refactor to use fd-relative") Pull Request: https://github.com/freebsd/freebsd-src/pull/2354
libifconfig: Add an SR-IOV VF status query Provide a public helper which retrieves, unpacks, and validates the versioned VF status nvlist. Validate the required VF indices and the shape and version of driver-specific extension namespaces while allowing unknown optional fields. The ioctl argument is not copied back when the command returns EFBIG. Start with a practical buffer and grow it geometrically rather than relying on the required length being observable. Use the helper in ifconfig so other consumers share the same transport and validation behavior.
rescue: Satisfy libifconfig's libnv dependency in crunched links libifconfig now calls nv(9) routines for the SR-IOV VF status query, so crunched builds that link the static library must also provide libnv. The per-program CRUNCH_LIBS_ifconfig hook cannot do this: crunchgen partially links per-program libraries into the program object and crunchide then localizes every symbol except the stub entry, so members absorbed there cannot satisfy references from another archive on the final link's library list. List libnv globally next to libifconfig.a in rescue(8) and bsdbox. This also makes the existing per-program libnv links redundant; remove them to avoid embedding private localized copies in the crunched binary. Fixes: https://cgit.freebsd.org/src/commit/?id=2d6114f6d26b ("libifconfig: Add an SR-IOV VF status query")
On powerpc64le with IEEE-128 long double, the long-double compiler-runtime helpers are the *kf* soft-float functions (built from the tf sources, renamed via -D in lib/libcompiler_rt/Makefile.inc) plus the complex multc3/__divtc3. They are compiled into libgcc_s.so by the powerpc64le SRCF block, but were never added to Symbol.map, so they stayed local and unexported. Every other IEEE-128 architecture already exports its scalar long-double runtime -- aarch64 and riscv list the tf helpers in GCC_4.6.0. powerpc64le was simply missed. Because the helpers are unexported, any clang-built shared library that uses long double leaves them undefined (permitted in a DSO), and linking an executable against that DSO then fails under lld's default --no-allow-shlib-undefined. For example science/harminv fails to link its binary against its own libharminv.so with undefined multc3/divtc3; at -O0, mulkf3/addkf3/__subkf3/__unordkf2 appear as well. Export the full runtime, gated on the PowerPC-specific LONG_DOUBLE_IEEE128 predefine so no other architecture is affected: complex multc3/divtc3 in GCC_4.0.0 (beside the other complex mul*c3), and the 28 scalar *kf* functions in GCC_7.0.0. Node placement follows glibc/gcc symbol-versioning history. Differential Revision: https://reviews.freebsd.org/D58248
Signed-off-by: Ishan Agrawal <iagrawal9990@gmail.com> Sponsored-by: Google LLC (GSoC 2026) Reviewed by: kp
Signed-off-by: Ishan Agrawal <iagrawal9990@gmail.com> Sponsored-by: Google LLC (GSoC 2026) Reviewed by: kp
We already verified that the attribute parser tables were correctly sorted. Now also verify that the command decoders are too. While here move the assertions into a constructor so we only run them once.
libusb: versioning symbols Reviewed by: bapt, kevans Sponsored by: The FreeBSD Foundation Differential Revision: https://reviews.freebsd.org/D55686
libusb: Mark defualt smybol tag as latest stable version As we might change the libusb ABI in 16, we should mark thje first version as FBSD_1.8 instead of 1.9. Since versioning patch has not landed for a long time, it makes sense to change it directly. Discussed with: kib Fixes: https://cgit.freebsd.org/src/commit/?id=527a82474cb3 ("libusb: versioning symbols") Sponsored by: The FreeBSD Foundation
Add BOOL_MAX and BITINT_MAXWIDTH macros for C23 compliance, and define the __STDC_VERSION_LIMITS_H__ feature test macro now that the header fully conforms to C23. Reviewed by: fuz Approved by: fuz (mentor) MFC after: 1 month Pull Request: https://github.com/freebsd/freebsd-src/pull/2352
Several standard library functions are specified to return an unqualified pointer while accepting a pointer to a potentially const-qualified object. N3020 addresses this behaviour, discarding qualifiers due to incompatible pointer types, by introducing qualifier-preserving macros for the affected set of standard library functions. Add `__qualsel()` helper to `<sys/cdefs.h>`, implemented using the generic selection, and define qualifier-preserving macros for that set of functions in `<string.h>`, `<wchar.h>`, and `<stdlib.h>`. Macros are gated on `_STDC_VERSION__ >= 202311L && !__cplusplus`, therefore there is no behavioural change for earlier C modes or C++ translation units. The kernel is likewise unaffected, as it does not include userland headers. As function-like macros, they are transparent except at a call site where the address-of operator is applied, the macro is suppressed via `#undef`, or the identifier appears in parenthesised form; all of which cause the underlying function designator to be used instead. Reviewed by: fuz Approved by: fuz (mentor) MFC after: 1 month Pull Request: https://github.com/freebsd/freebsd-src/pull/2288
In Linux these are maintained in separate places so a separate copy is needed, but in FreeBSD take advantage of the shared tree to avoid having a duplicate copy that can be stale. Reviewed by: np Sponsored by: Chelsio Communications Differential Revision: https://reviews.freebsd.org/D58575
Tables that have one element per protocol or address family were previously sized by AF_MAX + 1 since AF_MAX was off by one. Now that AF_MAX has been corrected, we need to apply the opposite correction to these tables. Fixes: https://cgit.freebsd.org/src/commit/?id=ddd850aa7720 ("sys/socket.h: Fix AF_MAX") MFC after: 3 days Sponsored by: Klara, Inc. Sponsored by: NetApp, Inc. Reviewed by: kevans Differential Revision: https://reviews.freebsd.org/D58827
FreeBSD's libusb has three components: libusb01, libusb10, and libusb20. libusb20 handles communication with character devices. We now requires a backend context for libusb20. The backend context contains contains the capsicumized usbctrl fd and usb directory (/dev/usb) fd so that the library user can enter the capiblity mode safely while using libusb. libusb10 is updated to support capabilities via a context option. Since libusb allows general read/write access, we preserve all possible capabilities when passing backend context to libusb20. It is the responsibility of the libusb user to call cap_enter() at an appropriate time. All base system tools using libusb and libusb20 have been updated to support Capsicum. Reviewed by: adrian, markj Sponsored by: The FreeBSD Foundation Differential Revision: https://reviews.freebsd.org/D51865
PR: https://bugs.freebsd.org/bugzilla/show_bug.cgi?id=296234(exp-run) Relnotes: yes Sponsored by: The FreeBSD Foundation Differential Revision: https://reviews.freebsd.org/D57772
Reviewed by: fuz Approved by: fuz (mentor) MFC after: 1 month Differential Revision: https://reviews.freebsd.org/D58842
Define the __STDC_VERSION_STDIO_H__ feature test macro now that the header fully conforms to C23. Reviewed by: fuz Approved by: fuz (mentor) MFC after: 1 month Differential Revision: https://reviews.freebsd.org/D58842
When fts_read() descends into a directory whose children were already prefetched by fts_children() (as ls -R does), it changed directory using p->fts_name instead of p->fts_accpath. With a trailing slash on a relative root path (e.g. 'dir/'), the bare name was resolved relative to the wrong directory, so every sibling directory after the first failed with ENOENT and was reported as FTS_DNR. This manifested as 'ls -lR dir/' skipping the contents of all but the first subdirectory. Restore the use of p->fts_accpath, matching the behavior prior to 4bd01d6ae016. Add a regression test that reproduces the exact conditions: fts_children() on each directory, FTS_PHYSICAL without FTS_NOCHDIR, and a trailing slash on the root path. Reported by: Michael Butler <imb@protected-networks.net> Reviewed by: asomers Fixes: https://cgit.freebsd.org/src/commit/?id=4bd01d6ae016 Sponsored by: Google LLC (GSoC 2026) Pull Request: https://github.com/freebsd/freebsd-src/pull/2372
fts_build() stores a dup'd file descriptor in each directory entry's fts_dirfd. When a traversal is abandoned before completion and fts_close() is called, the cleanup loop freed each pending entry with free() without first closing its fts_dirfd, leaking one descriptor per pending directory. Close fts_dirfd before freeing each entry in the cleanup loop, matching the handling already applied to the dummy parent entry after the loop. Add a regression test that descends a couple of levels, abandons the traversal, closes, and asserts the open descriptor count is unchanged. PR: https://bugs.freebsd.org/bugzilla/show_bug.cgi?id=297557 Reported by: asomers Fixes: https://cgit.freebsd.org/src/commit/?id=4bd01d6ae016 Sponsored by: Google LLC (GSoC 2026) Reviewed by: asomers Pull Request: https://github.com/freebsd/freebsd-src/pull/2373
Add fts_openat() as a new entry point for fts(3). When dirfd is AT_FDCWD the behaviour is identical to fts_open(). Passing a pre-opened directory fd allows fts traversal inside Capsicum capability mode where path-based operations are not permitted. Capability mode users should use fts_parent->fts_dirfd + fts_name with openat(2) to access files. Reviewed by: asomers Relnotes: yes Sponsored by: Google LLC (GSoC 2026) Pull Request: https://github.com/freebsd/freebsd-src/pull/2273
To avoid any sort of POLA violation, this commit restores old guards and defines the C23 feature test macros in addition to them. This is to close off whole class of possible breakage, rather than patching it case by case. Reported by: dim Reviewed by: dim, dteske, fuz Approved by: dim, dteske (mentor), fuz (mentor) MFC after: 1 month Differential Revision: https://reviews.freebsd.org/D58911
exterr: relax format restrictions Rather than passing the format string to printf and forcing the arguments to be (u)intmax_t, partially parse format strings and cast p1 and p2 to the correct type before running the individual format though printf. This restructure has a couple motivatations: - We can skip formats that make no sense (floating point, %n, etc.). - It is possible to special case the printing of pointers in the CHERI case. The first case is motivated by a suggestion from the audiance at one of Kirk's BSDCan talks on exterr to allow userspace to set exterr status. Allowing arbitrary format strings including %n creates a write-what-where gadget so we need to not do that. The second case is motivated by our experinces with CHERI and debugging mmap issues using a different textual error reporting framework. With CHERI, pointers are more than integer addresses and it's useful to include more details. Doing so will follow in a future commit. When the new code encounters an inappropriate format it includes a diagnostic and in most cases prints the format untouched. Reviewed by: kib Effort: CHERI upstreaming Sponsored by: Innovate UK Differential Revision: https://reviews.freebsd.org/D58058
exterr: Fix build with GCC on 32-bit architectures
Use an intermediate uintptr_t cast to avoid casting a uint64_t value
directly to void * on 32-bit platforms (including lib32 builds).
lib/libc/gen/uexterr_format.c: In function 'uexterr_format_msg':
lib/libc/gen/uexterr_format.c:248:35: error: cast to pointer from integer of different size [-Werror=int-to-pointer-cast]
248 | PFMT(fmt, (void *)ARG(nextarg));
| ^
lib/libc/gen/uexterr_format.c:144:47: note: in definition of macro 'PFMT'
144 | psz = snprintf(buf, bufsz, f, a); \
| ^
Reported by: GCC 15
Fixes: https://cgit.freebsd.org/src/commit/?id=2f024a7cfddd ("exterr: relax format restrictions")
The UEXTERROR(3) macro is a partial analog to EXTERROR(9) that sets the current user exterror state and errno. The main difference is that it returns no value and sets errno directly since that's the typical pattern in libraries. While here move the storage and constructor for single-threaded program's uexterr to its own file. Reviewed by: kib Effort: CHERI upstreaming Sponsored by: Innovate UK Differential Revision: https://reviews.freebsd.org/D58059
libusb upstream uses int for register handler. This causes some library user (like pyusb) to assume that we have int in all implementations and therefore provides a 4 byte storage only. This causes Segmentation fault as we will right the pointer. Reviewed by: adrian Sponsored by: The FreeBSD Foundation Differential Revision: https://reviews.freebsd.org/D54211 (cherry picked from commit ce9ced951a0b9d004a3b007d4ac6e9087a1301a2)
If next_callback_id wraps we could end up with two callbacks with the same ID. I recommitted the original change despite this issue in order to fix the libusb API as soon as possible after SHLIB_MAJOR was bumped in commit 527a82474cb3 (libusb: versioning symbols). It's very unlikely in practice that software will register and deregister a sufficient number of callbacks to trigger this, but it is a real issue to be fixed in a subsequent commit. Sponsored by: The FreeBSD Foundation
Instead, just reset next_callback_id to 1 at INT_MAX. The potential for duplicate callback IDs remains. Sponsored by: The FreeBSD Foundation
- Implement bsearch_s() as per §K.3.6.3.2 in C23, first specified in C11. It behaves identically to bsearch(), except the callback is called with a third argument, context, which is passed through from the caller, and it also performs runtime constraint checking on its arguments. - Document bsearch_b(), bsearch_s(), and add history section - Add rudimentary unit tests for bsearch(), bsearch_b(), and bsearch_s() Reviewed by: dteske, fuz Approved by: dteske (mentor), fuz (mentor) MFC after: 1 month Differential Revision: https://reviews.freebsd.org/D58876
libusb_hotplug_register_callback() runs the newly registered callback over the already-enumerated device list when LIBUSB_HOTPLUG_ENUMERATE is set. A hotplug callback returning non-zero means "deregister me", and the enumerate loop honours that by freeing the handle and setting it to NULL. Since commit 6bda9f26d2ed changed libusb_hotplug_callback_handle from a pointer to an int, the tail of the function unconditionally dereferences that handle, so any caller that passes LIBUSB_HOTPLUG_ENUMERATE, a non-NULL handle pointer, and a callback that returns non-zero on a matching device crashes inside libusb. This is a normal usage pattern and it was safe before the conversion, when the equivalent line simply stored NULL. Report the reserved id 0 instead. The allocator hands out ids starting at 1, and libusb_hotplug_deregister_callback() already ignores 0, so this restores the pre-conversion behaviour. Signed-off-by: yuvrajnode <yuvrajsinghrock1221@gmail.com> Reviewed by: aokblast Fixes: https://cgit.freebsd.org/src/commit/?id=6bda9f26d2ed ("libusb: change callback register handler to int") Pull Request: https://github.com/freebsd/freebsd-src/pull/2383 Closes: https://github.com/freebsd/freebsd-src/pull/2383
libusb_hotplug_register_callback() resolves its context with GET_CONTEXT() and then immediately reads ctx->no_discovery and ctx->usb_event_mode, but only checks "ctx == NULL" afterwards. GET_CONTEXT() falls back to usbi_default_context, which is NULL before libusb_init() and is reset to NULL by libusb_exit(). An application that calls libusb_hotplug_register_callback(NULL, ...) without an initialised default context therefore crashes on the ctx->no_discovery read, instead of getting the LIBUSB_ERROR_INVALID_PARAM the existing guard was clearly written to return. Move the argument validation ahead of the first dereference. None of the validated arguments depend on the context, so no other ordering constraint is affected. Signed-off-by: yuvrajnode <yuvrajsinghrock1221@gmail.com> Reviewed by: aokblast MFC after: 2 weeks Pull Request: https://github.com/freebsd/freebsd-src/pull/2384 Closes: https://github.com/freebsd/freebsd-src/pull/2384
libc: Remove incorrectly defined __STDC_VERSION_STDBOOL_H__ C23 does not define __STDC_VERSION_STDBOOL_H__ for <stdbool.h>. Feature test macros of this form only apply to headers where the standard explicitly mandates one. Reviewed by: fuz Approved by: fuz (mentor) MFC after: 1 month Differential Revision: https://reviews.freebsd.org/D59136
libc: Add <time.h> C23 feature test macro Define the __STDC_VERSION_TIME_H__ feature test macro as the header fully conforms to C23. Reviewed by: fuz Approved by: fuz (mentor) MFC after: 1 month Differential Revision: https://reviews.freebsd.org/D59134
libc: Add <setjmp.h> C23 feature test macro Define the __STDC_VERSION_SETJMP_H__ feature test macro as the header fully conforms to C23. Reviewed by: fuz Approved by: fuz (mentor) MFC after: 1 month Differential Revision: https://reviews.freebsd.org/D59135
libc: Fix C23 version macro visibility In headers that existed prior to C23, these should be visible only in C23 or BSD mode. Fixes: https://cgit.freebsd.org/src/commit/?id=0fe73dcf7c32 ("libc: Add <assert.h> C23 feature test macro") Fixes: https://cgit.freebsd.org/src/commit/?id=1f09e354297c ("sys/limits.h: Add BOOL_MAX, BITINT_MAXWIDTH, and C23 feature test macro") Fixes: https://cgit.freebsd.org/src/commit/?id=cd0727ec709b ("libc: Add <stdio.h> C23 feature test macro") Fixes: https://cgit.freebsd.org/src/commit/?id=fc9d02cb29ed ("libc: Add <time.h> C23 feature test macro") Fixes: https://cgit.freebsd.org/src/commit/?id=4aeed6e9d213 ("libc: Add <setjmp.h> C23 feature test macro") Reviewed by: fuz, kfv, dteske Differential Revision: https://reviews.freebsd.org/D59272
Sponsored by: Rubicon Communications, LLC ("Netgate")
Reviewed by: markj Sponsored by: The FreeBSD Foundation MFC after: 1 week Differential revision: https://reviews.freebsd.org/D58586
Fix nlist(3) consumers that either expected our toolchain to prepend an underscore to symbol names or expected nlist(3) to ignore the mismatch, as it did until we overhauled it back in May. While here, also fix cases where the last element in the list had an empty string instead of NULL as sentinel. MFC after: 3 days Fixes: https://cgit.freebsd.org/src/commit/?id=4617a6cb82a6 ("nlist: Handle multiple symbol tables") Reviewed by: kib, jhb Differential Revision: https://reviews.freebsd.org/D59254
lib9p was imported to add a 9p server to bhyve (and I believe this was the original motivation for writing it in the first place). Its external interfaces are kind of strange (from first-hand experience using it to implement an inetd-based 9p server) and undocumented. Moreover, upstream has been inactive for over five years. I suspect there are no third-party consumers. Let's make it a private library for now, so as to make it easier to rework external interfaces. If we get more code written against it, symbol versioning, and some documentation, we can revisit this decision. PR: https://bugs.freebsd.org/bugzilla/show_bug.cgi?id=297499 Reviewed by: jhb, emaste Differential Revision: https://reviews.freebsd.org/D58828
Reviewed by: kargl, kib Approved by: fuz (mentor) MFC after: 1 month Differential Revision: https://reviews.freebsd.org/D59288
Reported and tested by: bapt Reviewed by: bapt, markj Sponsored by: The FreeBSD Foundation MFC after: 1 week Differential revision: https://reviews.freebsd.org/D59348
Define __STDC_VERSION_INTTYPES_H__ now that the header fully conforms to C23. Reviewed by: fuz Approved by: fuz (mentor) MFC after: 1 month Differential Revision: https://reviews.freebsd.org/D59382
Define __WCHAR_WIDTH in sys/_types.h and derive WCHAR_WIDTH from that, the same way as WCHAR_MIN and WCHAR_MAX, in both <wchar.h> and <stdint.h> as per C23 §7.31.1 and §7.22.3.4, respectively. Reviewed by: fuz Approved by: fuz (mentor) MFC after: 1 month Differential Revision: https://reviews.freebsd.org/D59385
PR: https://bugs.freebsd.org/bugzilla/show_bug.cgi?id=290099 MFC after: 3 days Sponsored by: The FreeBSD Foundation
Request RTEXT_FILTER_VF through route Netlink and parse the common VF status schema into typed public structures. Preserve per-field presence using IFLAF_VF_* attribute numbers as mask bit indices so callers can distinguish omitted values from false or zero. Validate required VF indices, repeated driver namespaces, and their versioned typed fields while allowing unknown optional attributes. Return VF records through a pointer vector so append-only growth of the public VF structure does not change the array stride seen by existing consumers. Sponsored by: BBOX.io Differential Revision: https://reviews.freebsd.org/D58777
Utilize the ARCHLEVEL framework from libc/amd64/string to provide the way for runtime selection of the implementation, if wanted. Reviewed by: fuz, kfv Discussed with: kargl Sponsored by: The FreeBSD Foundation MFC after: 1 week Differential revision: https://reviews.freebsd.org/D59462
When doing a basic `jls -n`, jls(8) will jailparam_get() the mac.label for every jail on the system using the same set of jailparams, and thus the same jp_value. We only init the mac_t the first time, so the first jail would populate it with `?` from /etc/mac.conf and the resulting jail_get(2) would clobber it with the empty string, then a second jail would try to pass the empty string to the kernel and fail because it must have a non-zero length. Fix it by invoking jps_get() every time. Drop some comments to note that jps_get() will be invoked with zero || garbage from previous call, and be sure that we don't leak our previous mac_t. There aren't any other jps_get implementations at this time, so this shouldn't cause any unexpected problems. Reported by: ivy Reviewed by: jamie Differential Revision: https://reviews.freebsd.org/D57280
Reviewed by: fuz Approved by: fuz (mentor) MFC after: 1 week Differential Revision: https://reviews.freebsd.org/D59553
Replaced hardcoded jail parameter name strings in the base userland consumers with the JAIL_PARAM_* constants. No functional change. Reviewed by: jamie, adrian Differential Revision: https://reviews.freebsd.org/D59577
This fixes rounding at the last bit for subnormals. PR: https://bugs.freebsd.org/bugzilla/show_bug.cgi?id=298260 Reviewed by: kib MFC after: 1 week Differential revision: https://reviews.freebsd.org/D59579
PR: https://bugs.freebsd.org/bugzilla/show_bug.cgi?id=298260 Reviewed by: kib MFC after: 1 week Differential revision: https://reviews.freebsd.org/D59579
Event: EuroBSDcon 2026 Requested by: fuz Reviewed by: fuz, rmacklem Differential Revision: https://reviews.freebsd.org/D59588
Required to prevent function-like macros with the same name from being expanded in the definitions once they become active in a later C mode. Without the parentheses, the macro would rewrite the declarator, and the file would consequently fail to compile. This style is already used for similar cases such as mempcpy(). Reviewed by: fuz Approved by: fuz (mentor) MFC after: 1 week Differential Revision: https://reviews.freebsd.org/D59600
Reviewed by: dteske, fuz Approved by: dteske (mentor), fuz (mentor) MFC after: 1 month Differential Revision: https://reviews.freebsd.org/D59612
This was missed when bringing the kexec framework in. Fixes: https://cgit.freebsd.org/src/commit/?id=e02c57ff37e ("kern: Introduce kexec system feature (MI)") Reviewed by: kib Sponsored by: Hewlett Packard Enterprise LP Differential Revision: https://reviews.freebsd.org/D59696
sysconf(8) --version now prints the library version alongside its own so each can move on its own clock. Assigning a bitmask to bool already converts zero/nonzero; drop the redundant != 0 (fuz). Reviewed by: fuz, kfv Differential Revision: https://reviews.freebsd.org/D59720
Reviewed by: fuz Differential Revision: https://reviews.freebsd.org/D59745
db/hash: Harden hash(3) database code The hash(3) database code does not validate the on-disk database header, leaving it open to several OOB read and write vulnerabilities. This change adds basic header validation and array bounds checking to parts of the hash(3) code that can be manipulated by messing with the database header. Reviewed by: kevans Sponsored by: Klara, Inc. MFC after: 1 month Differential Revision: https://reviews.freebsd.org/D58822
libc/db: Correct a typo in the the hash tampering test Fixes the gcc build. Fixes: https://cgit.freebsd.org/src/commit/?id=7f5f07b139a5
Copy the descriptor into a buffer of at most 64 MiB (raise it with BSDCONF_MAX_BYTES) and tokenize with bsdconf_scan(), the walker bsdconf_put() already uses. Input above the cap fails with EFBIG. Bump libbsdconf to 1.2.0 and sysconf(8) to 2.0. Suggested by: fuz Reviewed by: fuz Differential Revision: https://reviews.freebsd.org/D59751
Approved by: kib Differential Revision: https://reviews.freebsd.org/D59929
Provide a tunable env variable to request immediate GC on destroy to allow to return to the previous behavior. Based on the report and patch by Atle Solbakken <atle.solbakken@gmail.com>. PR: https://bugs.freebsd.org/bugzilla/show_bug.cgi?id=268532 Sponsored by: The FreeBSD Foundation MFC after: 1 week
GCC 16[0][1] now requires a wrapper lib to handle `--push-state --as-needed -lgcc_s --pop-state` logic[1] for `gcc` invocations. `g++` will still unconditionally use `-lgcc_s`. This fixes buildworld with gcc16. [0] https://gcc.gnu.org/git/?p=gcc.git;a=commit;h=8a99fdb70493df1294b53406913e5ea1fc971c13 [1] https://gcc.gnu.org/bugzilla/show_bug.cgi?id=123396 Reviewed by: jhb MFC after: 3 days Sponsored by: The FreeBSD Foundation Differential Revision: https://reviews.freebsd.org/D59511
Reviewed by: jhibbits Approved by: olce (mentor) MFC after: 2 weeks Sponsored by: FreeBSD Foundation Differential Revision: https://reviews.freebsd.org/D59719
The regular powerpc64 core probes only checked the kernel ELF and thus also matched minidumps. In particular, the powerpc64le probe could claim a minidump before the minidump backend and then reject it as an invalid ELF core. Exclude minidumps from both regular powerpc64 probes and add a regression test that verifies a powerpc64le minidump reaches the minidump parser. Reviewed by: jhibbits Approved by: olce (mentor) Fixes: https://cgit.freebsd.org/src/commit/?id=f4eb39ba6bc9 ("[PowerPC64LE] libkvm powerpc64le support.") MFC after: 2 weeks Sponsored by: FreeBSD Foundation Differential Revision: https://reviews.freebsd.org/D59719
PowerPC64 radix minidumps use a 64KB root directory followed by three levels of 4KB page tables. Add an MMU backend which walks those tables, handles large-page leaves, and decodes their always-big-endian entries on both powerpc64 and powerpc64le. Reject old radix minidumps whose pmap section is empty with a specific diagnostic. Add a synthetic powerpc64le dump test which reads a page through both its kernel and direct-map addresses. PR: https://bugs.freebsd.org/bugzilla/show_bug.cgi?id=298532 Reviewed by: jhb Approved by: jhb (mentor) MFC after: 2 weeks Sponsored by: FreeBSD Foundation Differential Revision: https://reviews.freebsd.org/D59743
The kernel sends PFR_A_AF as a u32, although it is documented and parsed as a u8, and libpfctl decoded it with snl_attr_get_uint32() straight into the u8 pfra_af. That overwrote pfra_net, pfra_not and pfra_fback, and left pfra_af zero on big-endian hosts. Decode it via a temporary, and accept a u8 as well, so that the kernel can be corrected later without breaking libpfctl. Reviewed by: kp Approved by: kp (mentor) Fixes: https://cgit.freebsd.org/src/commit/?id=f27e44e2e3b5 ("pf: convert DIOCRGETADDRS to netlink") Sponsored by: Rubicon Communications, LLC ("Netgate") Differential Revision: https://reviews.freebsd.org/D60110
The chunked table address functions (set, add, del, clr_astats) add each chunk's result to the caller's counter without initialising it. pfctl reuses nadd for the number of tables created, so a replace that also creates the table is off by one: pfctl -t foo -T replace 192.0.2.1 reports "2 addresses added". Zero the counters first, as pfctl_test_addrs() already does. Remove the workaround for the add case from pfctl (da64f6e047b5), which is no longer needed. Add a regression test. Reviewed by: kp Approved by: kp (mentor) Fixes: https://cgit.freebsd.org/src/commit/?id=08ed87a4a276 ("pf: convert DIOCRSETADDRS to netlink") MFC after: 1 week Sponsored by: Rubicon Communications, LLC ("Netgate") Differential Revision: https://reviews.freebsd.org/D60174
Rename am_lock description from autofslk -> autfsm. The lock description, autofslk, is used as the description for autofs_softc->sc_lock, which is used to protect autofs requests and the like as opposed to am_lock which protects autofs nodes for a given mount. This change allows witness to distinguish different lock orders for each lock. Reviewed by: kib Differential Revision: https://reviews.freebsd.org/D57972
The OpenZFS merge 80aae8a3f8aa introduced HAVE_SIMD() which checks for HAVE_TOOLCHAIN_* defines via simd_config.h. The kernel module Makefile was updated, but kern.pre.mk (static kernel build) and the libzpool/libzfs Makefiles were missed, still using the old HAVE_SSE2 etc. names. This caused all vectorized raidz, fletcher, and blake3 implementations to be compiled out.
The consequences are: - for nfs exports and fhopen(2), unlinked but still referenced inodes are accessible - for ffs_vput_pair() with unlock_vp = false, spurious ESTALE is not returned when the inode is still alive but unlinked Note that tmpfs does not return ESTALE for the unlinked nodes. The same behavior is claimed for Linux in https://github.com/openzfs/zfs/issues/18699 Reviewed by: rmacklem Sponsored by: The FreeBSD Foundation MFC after: 1 week Differential revision: https://reviews.freebsd.org/D57982
Fix two locking violations that could happen during execve, while
executing a file stored on fusefs. Both would cause panics on an
INVARIANTS kernel after 15.0, or a DEBUG_VFS_LOCKS kernel prior to that.
Neither is likely to be noticeable on a release kernel.
* Don't assume that the vnode is exclusively locked during VOP_CLOSE.
It usually is thanks to !MNTK_LOOKUP_SHARED, but isn't during execve,
which locks the vnode outside of the lookup path.
* Totally rewrite fuse_io_invalbuf. It's had a number of problems ever
since its original introduction[^1]:
- Don't assume that the vnode is exclusively locked. That assumption
failed during execve just like the assumption in fuse_vnop_close.
- Don't livelock forever if vinvalbuf returns ENOSPC or EDQUOT.
- Don't attempt to handle multiple threads calling this function at
the same time. That would be impossible if the vnode truly were
exclusively locked. So the code was dead. Or it would've been, if
the assumption hadn't been wrong. Furthermore, both vinvalbuf and
vnode_pager_clean_sync only require a shared vnode lock, and are
already capable of dealing with multiple simultaneous callers.
- Using fvdat->flag in this way would require some sort of mutex
protection, if the vnode weren't exclusively locked.
* Add new test cases that trigger both of the aforementioned panics.
[^1]: https://github.com/glk/fuse-freebsd/commit/efe6eb3005e7633b4e31d5e453eacbaa0cba42fa
PR: https://bugs.freebsd.org/bugzilla/show_bug.cgi?id=295957
Reported by: dan.kotowski@a9development.com
MFC after: 2 weeks
Sponsored by: ConnectWise
Reviewed by: markj
Differential Revision: https://reviews.freebsd.org/D57536
Previously, an_vnode_lock was initialized with SX_NOWITNESS to silence
lock order reversals. The reversals would occur when autofs_node_vn()
was called with the directory vnode lock held, then lock an_vnode_lock,
then lock the vnode attached to the autofs node. It looked like:
directory vnode -> an_vnode_lock -> vnode attached to autofs node
The established lock order is now vnode -> an_vnode_lock
Currently, we don't have to worry about losing an autofs node during the
unlock/lock as autofs nodes are only removed during an unmount() after
vflush(). When autofs_node_vn() is called, the mountpoint has either
been busied (preventing unmount) or a directory vnode is locked which
prevents vflush() from finishing until the directory vnode is unlocked.
Reviewed by: kib
Differential Revision: https://reviews.freebsd.org/D57857
Commit 016570c4463d modified the client to handle the upgrade of a read delegation to a write delegation, where the server provides the same delegation stateid to the client. However, it failed to check if the delegation structure was currently in use. Without this patch, if the structure was in use, a use after free could occur. This patch handles the "in use" case by copying the necessary fields into the current/old structure and free's the new one instead of the old one that is "in use". PR: https://bugs.freebsd.org/bugzilla/show_bug.cgi?id=296224 MFC after: 2 weeks
Skip UFS2 fs_metaspace upper-bound validation that rejects NetBSD FFSv2 WAPBL filesystems due to differing superblock layouts. Detect the condition during mount instead and permit read-only mounts while rejecting read-write mounts with EROFS. This follows NetBSD's recommendation for systems without WAPBL support and avoids modifying unsupported journal metadata. PR: https://bugs.freebsd.org/bugzilla/show_bug.cgi?id=296022 Signed-off-by: Ricardo Branco <rbranco@suse.de> Reviewed by: imp, kirk Pull Request: https://github.com/freebsd/freebsd-src/pull/2279
nfsd: Garbage collect stray NFSv4 state When a file is deleted on the NFS server by another client, any NFSv4 state related to that file is left stranded. This happens because the NFSv4 operations that free the state use a CFH, which is set by a PutFH operation. However, the PutFH fails with ESTALE because the file has been deleted. This patch adds a function called nfsrv_freestrandedstate() that frees all the NFSv4 state related to a file and calls this function when PutFH will be replying ESTALE. While here, a helper function was defined to handle free'ng of the nfslockfile structure and replaces the two places where nearly identical code does this. Reported by: Richard Purdie <richard.purdie@linuxfoundation.org> Tested by: Michael Halstead <mhalstead@linuxfoundation.org> MFC after: 2 weeks
nfsd: Commit missing patches for c52bcd09c2a6 Oops, I missed the other files for the commit. This should fix the build. Pointy hat goes on me. MFC after: 2 weeks Fixes: https://cgit.freebsd.org/src/commit/?id=c52bcd09c2a6 ("nfsd: Garbage collect stray NFSv4 state")
Reviewed by: mckusick Discussed with: markj Tested by: pho Sponsored by: The FreeBSD Foundation MFC after: 1 week Differential revision: https://reviews.freebsd.org/D57658
Reviewed by: mckusick Discussed with: markj Tested by: pho Sponsored by: The FreeBSD Foundation MFC after: 1 week Differential revision: https://reviews.freebsd.org/D57658
After a bypassed VOP, nullfs mirrors the lower vnode's inotify state onto the upper vnode. The flags were checked with lockless reads before being updated with the asserting flag set/unset primitives, so two threads syncing the same vnode concurrently (or a sync racing a watch being established) could both decide to make the same change; the loser then trips the "flags already set" assertion on an INVARIANTS kernel. On other kernels the race is harmless. Keep the lockless check as the fast path, but re-make the decision under the vnode interlock before actually changing the flags. Reproduced in a 4-CPU VM with one thread cycling an inotify watch on a lower-filesystem file while several threads stat(2) the same file through a nullfs mount: the unpatched INVARIANTS kernel panics under this load, the patched kernel runs it to completion. Fixes: https://cgit.freebsd.org/src/commit/?id=f1f230439fa4 ("vfs: Initial revision of inotify") MFC after: 2 weeks Differential Revision: D58344 Reviewed by: markj Assisted-by: Claude Code (Fable 5)
Differential Revision: https://reviews.freebsd.org/D57898
If the server is closing (or the device node is going away), or if
devfs_set_cdevpriv() fails, cuse_client_open() returns with the server
reference taken at the top of the function still held and the newly
allocated client still linked on pcs->hcli. Since cuse_client_free()
has not been registered as the cdevpriv destructor at that point,
nothing ever undoes this work: every open() that races the is_closing
window permanently leaks one server reference and one cuse_client.
A leaked reference is fatal on server exit: cuse_server_free()
busy-waits in an uninterruptible pause("W", hz) loop until pcs->refs
drops to 1, which now never happens, so the exiting server process
(e.g. virtual_oss(8)) is left wedged in state "D", immune to SIGKILL,
cuse.ko is pinned (kldunload hangs too), and only a reboot recovers.
Before 634e578ac7b0 the is_closing error path dropped the reference by
calling devfs_clear_cdevpriv(), which ran the cuse_client_free()
destructor. That commit moved devfs_set_cdevpriv() after the
is_closing check to fix the panic paths, but left both error returns
without any cleanup.
Fix by calling cuse_client_free() directly on both error paths. The
client is fully constructed and linked on pcs->hcli at these points,
which is exactly the state cuse_client_free() expects.
PR: https://bugs.freebsd.org/bugzilla/show_bug.cgi?id=296291
Fixes: https://cgit.freebsd.org/src/commit/?id=634e578ac7b0 ("cuse: Fix cdevpriv bugs in cuse_client_open()")
Assisted-By: Claude Opus 4.8 (claude-opus-4-8)
Signed-off-by: giacomo <delleceste@gmail.com>
MFC after: 2 weeks
Reviewed by: christos
Pull-Request: https://github.com/freebsd/freebsd-src/pull/2324
Unlike RFC5661 (the original NFSv4.1 RFC), RFC8881 specifies that a NFS4ERR_DELAY reply to the SEQUENCE operation requires a reply using the same slot/sequence#. This patch fixes handling of this case, so it conforms to RFC8881. Reported by: J. David (j.david.lists@gmail.com) Tested by: J. David (j.david.lists@gmail.com) MFC after: 1 week
Commit 4d80d4913e79 added a check for nfsess_defunct already being set. This was incorrect because, once set, nfsess_defunct remains set and an additional recovery might be needed. This patch reverts this part of 4d80d4913e79. PR: https://bugs.freebsd.org/bugzilla/show_bug.cgi?id=297252 Suggested by: olivier MFC after: 3 days Fixes: https://cgit.freebsd.org/src/commit/?id=4d80d4913e79 ("nfs: Fix argument typo to avoid a crash")
Delegations in NFSv4.0 never worked well and, since the NFSv4.0 protocol is now deprecated, use of delegations for NFSv4.0 is disabled as far as the client can do so. It turns out that some Illumos NFSv4.0 server issues delegations anyhow (even when the callback path is specified as 0.0.0.0) and this can cause use after free problems. This patch deleted some cruft that did an nfsrpc_openrpc() call recursively when an NFSv4.0 server failed to issue a delegation when it had previously done so. This code was only meant to be an optimization and would have been rarely exercised. Since this recursive call of nfsrpc_openrpc() is in some of the backtraces in the bugzilla PR, getting rid of the cruft makes sense. It is not known if this helps w.r.t. the use after free problems at this time. PR: https://bugs.freebsd.org/bugzilla/show_bug.cgi?id=297233 MFC after: 3 days
This addresses a race when two vnodes attempt to call vfs_hash_insert(), but only one succeeds. Also, in case of an error from p9fs_reload_stats_dotl(), it marks the vnode for deletion. Reviewed by: kib MFC after: 1 week Differential revision: https://reviews.freebsd.org/D58632
Since autofs_lookup() calls into autofs_trigger_vn() to perform automounting, and autofs_trigger_vn() unlocks the vnode, it is possible for the unmount to start meantime. Then autofs_trigger() accesses freed memory. At this point, busy can be only done unblocking, and the transient failure must abort the trigger operation. This would cause spurious automounter errors, but at least should prevent accesses to the freed memory. PR: https://bugs.freebsd.org/bugzilla/show_bug.cgi?id=294361 Reviewed by: markj, rew Tested by: rew Sponsored by: The FreeBSD Foundation MFC after: 1 week Differential revision: https://reviews.freebsd.org/D58626
This is a waste of time and results in a use-after-free if linsysfs is loaded and a USB network interface is in use, since USB devices are disconnected at shutdown, which triggers a call into linsysfs, which then tries to destroy a pseudofs node which has already been purged. MFC after: 1 week Reviewed by: glebius Differential Revision: https://reviews.freebsd.org/D58359
Do the advisory aborts of the in-flight requests before flushing the vnodes. It should mostly eliminate the waits due to requests busying the mp. Reported and reviewed by: rew Sponsored by: The FreeBSD Foundation MFC after: 1 week Differential revision: https://reviews.freebsd.org/D58637
Previously fts_build() called _dup(_dirfd(dirp)) for every child entry, holding N simultaneous fds for a directory with N children. Redefine fts_dirfd: instead of a fd for the entry's parent directory, it is now a fd for the entry itself, set only for directory entries. One dup per directory in fts_build() instead of one per child. Close fts_dirfd during the directory post-order visit, before advancing to its sibling. To access a file using fd-relative operations, callers should use openat(ent->fts_parent->fts_dirfd, ent->fts_name, ...) instead of openat(ent->fts_dirfd, ent->fts_name, ...). The fd is valid until the directory's post-order visit (FTS_DP). Reported by: Mark Johnston <markj@FreeBSD.org> Fixes: https://cgit.freebsd.org/src/commit/?id=4bd01d6ae016 (fts: refactor to use fd-relative operations) Sponsored by: Google LLC (GSoC 2026) Reviewed by: asomers Pull Request: https://github.com/freebsd/freebsd-src/pull/2360
The backing file is already opened successfully, so there is no need to override the permissions. Reviewed by: des MFC after: 1 week Sponsored by: The FreeBSD Foundation Differential Revision: https://reviews.freebsd.org/D58832
Reviewed by: des MFC after: 1 week Sponsored by: The FreeBSD Foundation Differential Revision: https://reviews.freebsd.org/D58833
As of commit 9f5c4ef328 ("dounmount(9): temporarily enable recursion
for the covered vnode lock"), the unmount path handles recursion
automatically, so there's no longer a need to handle this case
in unionfs-specific code.
Reviewed by: kib, markj
Tested by: pho
Differential Revision: https://reviews.freebsd.org/D58858
rc+devd: Add growfs_postboot
In VM and cloud environments it is often possible to enlarge virtual
disks; this can be useful, for example, if a system is launched with a
small root disk and it later becomes clear that more space is needed.
On kernels which support run-time resizing of disks (for NVMe, this was
added in November 2025; some other disk types have supported this for
longer) a SIZECHANGE notification is sent to userland via devd.
Add a "nostart" rc.d script (runnable manually but not automatically at
boot time) and a devd script which invokes it when a notification
arrives. The rc.d script enlarges the "final partition" on partitioned
geoms, or the UFS filesystem or zpool device when triggered on a disk
containing either of those.
Reviewed by: imp, ziaee
MFC after: 2 weeks
Relnotes: Disk partitions and filesystems can be enlarged
automatically when disks grow by setting
growfs_postboot_enable=YES in /etc/rc.conf.
Sponsored by: Amazon
Differential Revision: https://reviews.freebsd.org/D58582
growfs_postboot: Invoke via service(8) Use service(8) rather than invoking /etc/rc.d/growfs_postboot directly. Requested by: bapt Reviewed by: bapt MFC after: 2 weeks Fixes: https://cgit.freebsd.org/src/commit/?id=5a31987d4c39 ("rc+devd: Add growfs_postboot") Differential Revision: https://reviews.freebsd.org/D59145
Notable upstream pull request merges:
#16761 eb1738bcb FreeBSD: Enable Direct IO by default
#16747 -multiple On-demand log-spacemap flush
#18474 0d0eae2ab implement thorough scrub support (zpool scrub -t)
#18562 6721ab981 Calling thread IO
#18657 7ed1268d3 Fix race between device removal completion and pool export
#18707 0d1f3b1c6 RAIDZ: Fix parity regeneration/check condition
#18713 -multiple libzfs: fix key unload failure when unmounting an
encryption root
#18714 d902eec64 Disconnect metaslab tracing from default builds
#18716 e78fa488a zstream: multithreading
#18718 6acb99cb1 Do not return ESTALE for open-unlinked files
#18720 f607ef7e7 Fix insufficient locking in dedup verify
#18722 -multiple snapdir: misc cleanups
#18724 9bf75b4b1 Fix reads for blocks freed after being cloned
#18725 9b7642df9 Harden recv record validation
#18732 37f066e3c FreeBSD: Wire sha512 offload to the build
#18736 -multiple Remove idmap/userns concept from core & FreeBSD platform
code
#18742 d63f14057 libzutil: keep valid spare and l2cache paths on import
#18749 f217627d4 Fix receive of split large blocks with a short trailing
chunk
#18754 83bf40784 Use a single creation time for a recursive snapshot
#18755 ebd9e2765 zstream: add "drop records" chain module
#18763 7d565cdfa zfs bookmark: add recursive (-r) bookmark creation
#18765 50557cc80 zpool initialize: add -z to write zeroes
#18766 0467ba01c Add SECURITY.md policy filE
#18773 ffde77051 Fix deduce_nblkptr for large dnodes on receive
#18784 -multiple libzfs: clarify the raw incremental "IV set guid mismatch"
error
#18793 9a02e53ad libzfs: don't abort receiving a raw encrypted send that
carries holds
#18795 30426217b Rate limit Direct I/O verify zevents
#18798 73b202451 Bound explicit user prefetch to a fraction of the ARC
#18801 2a5331f8e Fix dmu_zfetch_prime() assuming a stream was created
#18802 fb5fdefd8 libzfs: fallback VDEV_UPATH to VDEV_PATH for non-DM
devices
#18809 bd2d87399 zdb: output refcounts from verify_spacemap_refcounts()
#18811 5d530b8be ABD: Validate borrowed buffer length
#18812 5d687a975 dbuf: use dirty record size for overridden writes
#18819 -multiple ddt: fix refcount bypass and gang member leak for dedup
gang blocks
#18821 87e2e4047 BRT: Implement partial bv_entcount writing
#18822 -multiple dnode_sync: Relax constraint on indirect freeing - #18822
#18827 d98fa72ca L2ARC: bound the rebuild by the write hand on a first
sweep
#18831 90cdd9aef Batch object reallocation syncs in zfs receive
#18833 466d90e15 zed: let autoexpand see capacity changes on partitioned
disks
#18835 1ac3f2786 mmp: skip non-writeable vdevs during activity check
#18838 a04c40138 DDT: Fix several bugs in pruning
#18841 0bb41751f zpool export: return EBUSY when zvol minors are in use
#18848 3020c18ca dmu_recv: Avoid potential null deref
#18855 8cdd9b2b7 mmp: do not require writes to mirror legs the config
marks absent
#18858 596c7a0ec Reduce dp_lock scope for MOS writes
#18859 3ebf3ffc4 Predict and throttle buffers dirtied by the sync context
#18860 -multiple Further parallelize block cloning
#18865 412b17a29 Fix DMU bonus hold leak on I/O error
#18867 b5bb1f816 Add missing checks to zfs_clone_range_replay()
#18868 3df5b2e7f libzfs: String trimming should not operate out of bounds
#18869 be62c5385 libzfs: Do not call munmap() when mmap() fails
#18871 74e76f2d8 libzfs: don't truncate a resolved vdev path in
zpool_vdev_name()
#18874 7312322b5 nvpair: Fix operator precedence
#18883 3bd8cefdc libzfs: don't read a dataset handle after closing it in
resume send
#18886 b4f70cb9e Fix negative time overflows in DDT pruning
#18888 674b1ae5c DDT: Make ddt_zap_walk() use cursor _by_dnode functions
#18889 fd1ae59c7 DDT: Skip DDT log lookups in ddt_prune_walk()
#18892 -multiple zhack: add "mmp reclaim" to recover a pool stranded by MMP
#18897 216de07d8 arc: fix race between arc_release() and arc_read_done()
#18899 28afe8b08 CodeQL: Flag implicit compare-then-assign in branch
conditions
#18917 0a79c039f arc: save on buf_hash_find() call at arc_read_done()
#18920 7026d3335 Fix race condition in raidz expansion startup
#18927 be55e01bc FreeBSD: Do not leak TSD zfs_geom_probe_vdev_key on errors
#18928 5bd7d4466 scan: count skipped blocks as examined
#18937 a5d678864 Allow pool import with corrupted spare/l2arc configs
#18960 -multiple vdev_open: use calling credential to check for device
access
#18961 acbcdd92c dmu_recv_begin_check: dsl_dataset_rele() should be called
on ds
#18962 df3ff37fc dmu_redact_snap: Do proper cleanup on ENAMETOOLONG
#18964 84aa7e7e0 libspl: consult ZFS_HOSTID on FreeBSD as well
Obtained from: OpenZFS
OpenZFS commit: 84aa7e7e09f6a4ddad9ec40dbe9498d50184ed07
In the world of containers, mounting a unix(4) socket is a common practice to allow communication between processes within containers. For example, both Podman and Docker can expose a unix(4) socket, and that same unix(4) socket can be mounted as a file accessible to a process inside a container, allowing that application to control Podman or Docker. Another example is PHP-FPM with NGINX, where, instead of using TCP/IP for communication between containers, a unix(4) socket is sufficient. However, nullfs(4) and all related components do not allow mounting a VSOCK on top of another. The current workaround involves creating the socket in a directory and mounting that directory. This is an option, though it does not provide a good user experience compared to directly mounting a VSOCK on top of another, since the application that creates the socket may create other sockets in that directory, and the user may not wish to share them, or, worse yet, applications that create unix(4) sockets may not provide any authentication at all, as they may assume that security at the file system level is sufficient. Reviewed by: dfr@ Approved by: dfr@ Relnotes: yes Differential Revision: https://reviews.freebsd.org/D59158
This patch adds assorted bits needed by the nfsclrdma.ko module that implements client side NFS over RDMA. With this commit, the glue required by the nfsclrdma.ko module is complete and it should load ok. It should not affect non-RDMA operation. I've specified a long MFC, since the module still requires extensive testing and, hopefully, a review. MFC after: 3 months
MFC after: 3 months Fixes: https://cgit.freebsd.org/src/commit/?id=884ee8d6c9b4 ("nfscl: Add some glue for client side NFS over RDMA")
The OFED code checks for a vnet argument, but it is is not defined. Reported by: glebius MFC after: 3 months Fixes: https://cgit.freebsd.org/src/commit/?id=884ee8d6c9b4 ("nfscl: Add some glue for client side NFS over RDMA")
Linux resolves a path below /proc/self/fd/N in the directory the
descriptor names, and linprocfs makes /proc/<pid>/fd a symlink to
/dev/fd. Under linrdlnk the fdescfs node carries only VV_READLINK,
which namei will not walk through, so such a path fails with ENOTDIR.
Return the underlying vnode from fdesc_lookup for a non-final component,
or a trailing slash, reusing the machinery the nodup option already
uses. The last component is untouched, so open("/dev/fd/N") keeps its
dup(2) semantic; a descriptor with no vnode behind it, such as a pipe,
yields ENOTDIR.
Add ATF coverage for traversal, descriptor reuse, and preservation of
last-component and mount-option semantics.
Approved by: adrian (mentor)
Reviewed by: kib, adrian
Differential Revision: https://reviews.freebsd.org/D59393
Signed-off-by: Nick Price <nprice@FreeBSD.org>
Notable upstream pull request merges:
#18808 6d4ff3b4f Pool split leaks DTL spacemap objects
#18844 0833cfd64 Decline Direct I/O reads on a file handle after a benign
verify failure
#18853 39764eb06 Fix double free when cloning a block with a pending free
#18901 51fe1fb84 dmu_send: a meta-dnode range does not start where an
object does
#18903 904d4327f dmu_objset_open_impl: unregister prop callbacks on error
#18906 e45640415 zdb: add detailed diagnostics for MOS leaks and spacemap
refs
#18909 1b7143557 Keep a grown vdev from adopting an older pool's labels
#18912 -multiple Defer destruction of a snapshot that a mount is holding
open
#18921 33ae06bb6 ZIO: Batch vdev children completions
#18940 7853c27eb ZIO: Batch lightweight ZIOs
#18967 -multiple Speed up pool rewind and allow it to tolerate minor damage
#18968 b0fff48dc arc: harness buf_hdr's anon state invariant checks
#18969 b0da5c4ce dsl_crypt: validate every key change-key will rewrap
#18974 b19fa7c97 log_spacemap: make zfs_log_sm_blksz tunable
#18982 c878ee452 zstream: track and limit memory use
#18991 a1261384a zstream: self-tuning queues
#18992 fd639b075 zcp: support bookmarks in zfs.sync.destroy
#18993 0498c554d vdev_disk_open: fix ENOENT retry & cleanup error handling
#19003 5280a08f6 Write uberblocks only to vdevs used by the txg
#19008 c47591d5d arc: apply arc_min tunables set before arc_init() has run
#19016 cd578e518 dnode: fix heap-use-after-free in dnode_rele_and_unlock()
#19017 b824f3116 zstd: declare __asan_*_memory_region() for user space
builds
#19029 6860d1c2f DDT: Use proper size when calling kmem_free on
ddt_prune_entry_t
#19041 aa26ca67b Do not complete a device removal that hit IO errors
#19042 8e5bb1d4b Wait for the initialize and trim threads to exit
#19051 ba524e0d7 zpool: things get messy with 0 columns
#19053 69ef525a4 dsl_deadlist_space_range: use correct tree when summing
up space
#19059 0ee8ea01b FreeBSD: use sys/abi_compat.h for time32_t
#19062 e08bc2973 Skip interior dnode slots when receiving DRR_FREEOBJECTS
#19066 92a3904af Wait for the txg that finished a scan to sync
#19071 db07a5e0a dbuf: account dirty data for embedded writes
#19072 1a8e51491 Fix self-deadlock when cloning a range within the same
file
#19075 75bd31c28 L2ARC: do not feed a device while its rebuild is pending
#19083 56e8d234b Stop the TRIM and initialize threads when a vdev goes
offline
#19090 ab994acdd dmu: remove dmu_prealloc()
#19092 61fbdf1a2 zfs: spa_sync_upgrades() should only take lock when needed
Obtained from: OpenZFS
OpenZFS commit: 02b5baf13cbfdf9abd33b939a082b448951d9174
Without this patch, nd_xprt is only set for NFSv4.1/4.2. The nfsrdma.ko needs it to be set for all versions of NFS, so this one line patch does that. No semantics change for non-RDMA NFS service. MFC after: 3 months Fixes: https://cgit.freebsd.org/src/commit/?id=7144a1d58c5c ("nfsd: Add glue for the nfsrdma.ko module")
If a daemon never responds to FUSE_INIT, but some process attempts to access the mountpoint, and a different process (possibly the daemon itself) attempts to unmount, a deadlock would result. Fix this bug by blocking any thread that enters fuse_vfsop_root until the daemon responds to FUSE_INIT or it times out. That will block any thread attempting to lookup a path in the fuse mountpoint, before it even gets to fuse VOPs. Remove less thorough initialization checks in fuse_ticket_fetch and fuse_vnop_access that are no longer necessary. And add a test case. PR: https://bugs.freebsd.org/bugzilla/show_bug.cgi?id=287431 MFC after: 2 weeks Sponsored by: ConnectWise Reviewed by: js Differential Revision: https://reviews.freebsd.org/D59737
As result, we lock the devvp vnode around calls to VOP_FSYNC() on unmount. For instance, the vn_fsync_buf() implementation of fsync() needs exclusive lock on the vnode to guarantee that all dirty buffers are indeed synced. Reviewed by: markj Tested by: pho Sponsored by: The FreeBSD Foundation MFC after: 1 week Differential revision: https://reviews.freebsd.org/D59932
Notable upstream pull request merges:
#17864 e903655c5 zpool: Add zpool status -vv error ranges
#18820 -multiple zdb: account pending DDT-log frees in leak detection
#18884 -multiple zio_crypt: establish platform interface; rework common
code to use it
#19010 -multiple zfs_namecheck: reject '.' and '..' before a snapshot or
bookmark
#19031 994fb1703 Fix metaslab count assertion in metaslab_group_alloc()
for small vdevs
#19093 f5b2fc8e2 zfs_ctldir: make .zfs/snapshot/<name> btime the snapshot
creation time
#19097 2dece2a34 zstream: report invalid record context without assertions
#19102 78f49e1dd vdev_disk: simplify alignment checks for linear ABDs
#19108 -multiple Fix permanent errors misfiled into the scrub error log
#19110 fa4bc4dec spa_errlog: don't let one unresolvable entry hide the
whole error log
#19116 b6dde8a17 zio_crypt: free the key unwrap uios when decryption fails
#19122 -multiple Fix dd_lock contention on space accounting
#19126 c792317f2 Avoid dp_lock in the dirty space accounting
#19131 7fd87803e Make the spa_config_lock reader fast path lockless
#19137 a83d96835 libzfs: do not iterate snapshots for a mountpoint
changelist
#19152 ccce78729 ZIO: Order a batch arrival against the hold it releases
#19155 9d7260c9b Propagate vdev state change on probe failure
#19157 946be3a33 FreeBSD: remove legacy ioctl support
#19165 1d9c32c16 zpool: accept more redundant special and dedup vdevs
#19167 638abd22a zfs_domount: fix vfs_t double-free on root setup failure
Obtained from: OpenZFS
OpenZFS commit: 1f380a4f34da6068d22960807f53cda68a356add
When an fdescfs mount has the nodup option set, fdesc_lookup(/dev/fd/n) returns the vnode referenced by file descriptor n, rather than returning an fdescfs vnode. This meant that fd metadata attached to fd n was not preserved when reopening the file, which is contrary to the expected semantics for capsicum rights and the UF_RESOLVE_BENEATH fd flag. For regular fdescfs mounts, this metadata is copied via dupfdopen(). Fix the problem by passing up this metadata through the nameidata structure. Thus, if one opens /dev/fd/n, the returned fd will inherit UF_RESOLVE_BENEATH and the capability rights of fd n. Add some regression tests as well. Approved by: so Security: FreeBSD-SA-26:66.jail Security: CVE-2026-101304 Reported by: Jan Bramkamp Reviewed by: kib Sponsored by: The FreeBSD Foundation Differential Revision: https://reviews.freebsd.org/D59886
When ufs_rename() moves a directory to a new parent and ufs_dirrewrite() fails to rewrite its ".." entry, it reports "bad dir ... rename: missing .. entry" whatever the error. ufs_dirrewrite() never finds a missing "..", though: it fails with EIDRM when the ".." entry names another inode than the expected one, and otherwise only when the directory block cannot be read or written. The latter happens for every directory rename in flight when the device goes away under a forcibly unmounted file system, and the log then fills with reports of directories that are intact on disk. Report the EIDRM case with ufs_dirbad() in ufs_dirrewrite() itself, so that all of its callers get the same diagnostic, and drop the report from ufs_rename(). Errors from the lower layers are not reported, as usual for an I/O initiator. Reviewed by: kib MFC after: 2 weeks Differential Revision: https://reviews.freebsd.org/D60137
cuse: Fix hang on readv(2) and writev(2) with multiple iovecs uiomove() leaves an iovec it has just emptied as the current one, so cuse_client_read() and cuse_client_write() picked it up again on the next iteration, sent the server a zero-length command, and got zero bytes back. That left the residual count unchanged, so the loop never terminated and the call never returned. Step past empty iovecs at the start of every iteration. This also covers caller-supplied zero-length iovecs, which hung in the same way PR: https://bugs.freebsd.org/bugzilla/show_bug.cgi?id=293489 MFC after: 1 week Sponsored by: The FreeBSD Foundation Reviewed by: kib, markj Differential Revision: https://reviews.freebsd.org/D59822
cuse: Actually use cuse_modevent() We can call cuse_kern_init()/cuse_kern_uninit() here, rather than using SYSINIT/SYSUNINIT. MFC after: 1 week Sponsored by: The FreeBSD Foundation Reviewed by: kib Differential Revision: https://reviews.freebsd.org/D59862
cuse: Create /dev/cuse with MAKEDEV_CHECKNAME Since we now use make_dev_credf(), make sure to fail kldload if it returned NULL. MFC after: 1 week Sponsored by: The FreeBSD Foundation Reviewed by: kib Differential Revision: https://reviews.freebsd.org/D59863
cuse: Improve server cleanup Move cuse_server_unref()'s device cleanup loop into a new cuse_server_free_devs_locked(), and call it from cuse_server_free() instead. The cdevpriv destructor now destroys the server's devices before dropping its reference, which closes the clients using them, so that the destructor is always the one that takes the last reference. By the time cuse_server_unref() frees the server, the device list should be empty, so assert this. In cuse_kern_uninit(), delete the infinite loop which waits for all open /dev/cuse instances to exit, and instead call destroy_dev() directly, which runs their cdevpriv destructor. MFC after: 1 week Sponsored by: The FreeBSD Foundation Reviewed by: kib Differential Revision: https://reviews.freebsd.org/D59872
cuse: Remove unnecessary semicolon in cuse_convert_error() No functional change intended. MFC after: 1 week Sponsored by: The FreeBSD Foundation
cuse: Retire unnecessary CUSE_VERSION No functional change intended. MFC after: 1 week Sponsored by: The FreeBSD Foundation
cuse: Use make_dev_s() to create client devices make_dev_s() sets si_drv1 before the node is published in devfs, which avoids a race where cuse_client_open() could see it as NULL. It also now reports finer-grained errors on failure, instead of only ENOMEM. While here, drop the NULL checks on kern_dev in cuse_server_free_dev(), since a device is only added to the server's list once it has been created. MFC after: 1 week Sponsored by: The FreeBSD Foundation Reviewed by: kib Differential Revision: https://reviews.freebsd.org/D59874
cuse: Implement hot-unload cuse_kern_uninit() can hang on destroy_dev(), because of threads sleeping in CUSE_IOCTL_GET_COMMAND, so implement d_purge to wake them up before calling destroy_dev(). Also do not allow threads to go back to sleep if the is_closing flag has been set. MFC after: 1 week Sponsored by: The FreeBSD Foundation Reviewed by: kib Differential Revision: https://reviews.freebsd.org/D60022
cuse: Assert the server refcount Assert that the refcount does not underflow before decrementing it, and that it really is zero by the time the server is freed. MFC after: 1 week Sponsored by: The FreeBSD Foundation Reviewed by: kib Differential Revision: https://reviews.freebsd.org/D60043
cuse: Rename cuse_server_free() to cuse_server_dtor() This name is clearer, given that this function is the cdevpriv destructor callback. MFC after: 1 week Sponsored by: The FreeBSD Foundation
Kernel stuff (other than networking, filesystems, and drivers).
Move the check out of ktls_enable_(rx|tx) and into ktls_create_session. Reviewed by: gallatin, markj Sponsored by: Chelsio Communications Differential Revision: https://reviews.freebsd.org/D57973
TLS receive offload is really only beneficial for in-kernel use cases (such as NFS over TLS) or when using a hardware offload. In addition, several recent SAs have involved the TLS receive path, but the only current mitigation for those is to disable TLS offload entirely. Reviewed by: ziaee, gallatin, markj Relnotes: yes Sponsored by: Netflix Sponsored by: Chelsio Communications Co-authored-by: John Baldwin <jhb@FreeBSD.org> Differential Revision: https://reviews.freebsd.org/D57974
linuxulator: Fix O_PATH file descriptors errno for f*xattr(2) LTP open13 expects these operations to fail with EBADF, matching Linux behavior, but FreeBSD currently returns EOPNOTSUPP for fgetxattr() on an O_PATH fd Look up Linux fd-based xattr descriptors with getvnode() and route the operations through shared kern_extattr_*_fp() helpers so the O_PATH check and the extattr operation use the same referenced file. Apply the same EBADF handling to fsetxattr(), fremovexattr(), and flistxattr() so the xattr paths stay consistent. Signed-off-by: YAO, Xin <mr.yaoxin@outlook.com> PR: https://bugs.freebsd.org/bugzilla/show_bug.cgi?id=295537 Reviewed by: kib Pull Request: https://github.com/freebsd/freebsd-src/pull/2263
linuxulator: Fix operator precedence for LINUX_XATTR_FLAGS in setxattr() The LINUX_XATTR_FLAGS macro expands to (LINUX_XATTR_CREATE|LINUX_XATTR_REPLACE). Without parentheses around the macro expansion, the bitwise & operator has higher precedence than |, causing incorrect flag evaluation and a compiler warning. Add the missing parentheses around LINUX_XATTR_FLAGS to ensure correct operator grouping, matching the existing usage in getxattr(). Signed-off-by: YAO, Xin <mr.yaoxin@outlook.com> Fixes: https://cgit.freebsd.org/src/commit/?id=2c905456312b ("linuxulator: Fix O_PATH file descriptors errno for f*xattr(2)") Reviewed by: kib Pull Request: https://github.com/freebsd/freebsd-src/pull/2306
Move the atomic size-probe-and-read logic into a new linux_extattr_get_vp() function in linux_xattr.c instead of modifying the generic extattr_get_vp() in vfs_extattr.c. This keeps Linux-specific getxattr semantics (ERANGE on too-small buffer, EOPNOTSUPP to ENOATTR mapping) self-contained within the linuxulator. The function probes the attribute size and reads the data under a single vnode lock, preventing a TOCTOU race between the size probe and data read. Signed-off-by: YAO, Xin <mr.yaoxin@outlook.com> Reviewed by: kib Pull Request: https://github.com/freebsd/freebsd-src/pull/2263
CHERI: declare mem{cpy,move}_data
Declare kernel-only, provenance-discarding memcpy_data, and memmove_data
APIs intended to copy raw data which does not contain pointers (e.g.,
buffers on their way to or from network or storage devices). On CHERI
architectures, they will explicitly remove tags from capabilities,
removing any provenance. This reduces the risk of accidental spread of
pointers on CHERI systems.
Document that bcopy preserves pointer provenance.
Reviewed by: ziaee, kib, adrian, markj
Effort: CHERI upstreaming
Sponsored by: DARPA, AFRL, Innovate UK
Differential Revision: https://reviews.freebsd.org/D57662
CHERI: add sooptcopyinptr to preserve pointer provenance Most socket options don't involve pointers so make the default sooptcopyin discard provenance and add a sooptcopyinptr that preserves. Reviewed by: markj, emaste Effort: CHERI upstreaming Sponsored by: DARPA, AFRL, Innovate UK Differential Revision: https://reviews.freebsd.org/D57665
CHERI: make mem{cpy,move}(9) CHERI compatible
- Use intptr_t in place of long as the word type in the core copying
loop where aligned words a copied. This preserved the provenance of
any copied pointers.
- When working with the address of src or dst use ptraddr_t rather than
uintptr_t. This avoid ambigious provenance in expressions involving
multiple addresses.
As a minor tweak, rename the function to memmove since that is the
interface it implements (overlapping src and dst are permitted) and make
memcpy the alias rather than the other way around.
Reviewed by: kib, markj
Effort: CHERI upstreaming
Sponsored by: Innovate UK
Differential Revision: https://reviews.freebsd.org/D57965
fetch.9: fix a typo Fixes: https://cgit.freebsd.org/src/commit/?id=a1c52e05f571 ("CHERI: declare fueptr and suptr") Effort: CHERI upstreaming Sponsored by: Innovate UK
Fixes: https://cgit.freebsd.org/src/commit/?id=d15792780760 ("unix: new implementation of unix/stream & unix/seqpacket") Reviewed by: glebius MFC after: 1 week Differential Revision: https://reviews.freebsd.org/D57967
Attaching to a jail changes its root directory and its process credentials. These operations both require unlocking the jail, and also need allprison_lock unlocked. That means that if two threads are trying to attach to different jails at the same time, it's possible for the process to end up with one jail's root directory but the other jail's credentials. Solve this by forcing the process into single-threaded mode during system calls that attach to a jail (jail_attach, jail_attach_jd, and sometimes jail_set). Reviewed by: kib, markj MFC after: 3 days Differential Revision: https://reviews.freebsd.org/D57858
If git is installed and .git exists but git rev-parse failed to report a hash we previously produced just "-dirty" as the git revision. Gate the git commit count and -dirty check on the rev-parse passing. Reviewed by: jlduran Sponsored by: The FreeBSD Foundation Differential Revision: https://reviews.freebsd.org/D57995
It is a handy shortcut that will be used extensively in hwpstate_intel(4) and hwpstate_amd(4). Warn users that it panics if the parent bus does not provide the CPU_IVAR_PCPU instance variable. That condition should be tested by callers (doing so once is enough). Suggest to do that in driver's attach method. Reviewed by: jhb (code) Event: Halifax Hackathon 202606 Location: Seat 36K in AC667, waiting for a gate at Montréal-Trudeau Sponsored by: The FreeBSD Foundation Differential Revision: https://reviews.freebsd.org/D57897
This makes the header more self-contained. The symbol is needed only on 32bit arches, but the include file is provided unconditionally to make the namespace population predictable. PR: https://bugs.freebsd.org/bugzilla/show_bug.cgi?id=296489 Sponsored by: The FreeBSD Foundation MFC after: 1 week
When a GID table entry is empty or not yet present in the cache, show_port_gid() falls back to printing a zero GID. Use the existing GID_PRINT_FMT/GID_PRINT_ARGS helpers instead of Linux's %pI6 format, which FreeBSD printf treats as a pointer followed by "I6". This makes empty GID sysctl entries consistently report 0000:0000:0000:0000:0000:0000:0000:0000. Tested by: Wafa Hamzah <wafah@nvidia.com> (mlx5_ib) Reviewed by: jhb, kib Sponsored by: NVIDIA Networking Fixes: https://cgit.freebsd.org/src/commit/?id=6a75471dbcf0 ("OFED: Various changes from Linux 4.19") Differential Revision: https://reviews.freebsd.org/D58042
inotify: Unconditionally generate IN_IGNORED events for files/dirs The implementation previously only generated an IN_IGNORED event for a deleted watched file if the watch explicitly requested IN_DELETE_SELF. This is not correct, IN_IGNORED should always be raised when the watched subject is deleted. Adjust the implementation of inotify_log_one() accordingly. This also fixes a problem where a deleted watched file's watch would not be removed if IN_DELETE_SELF was not in the watch's event mask, in which case the unlinked vnode would linger until the inotify descriptor itself is closed. Add a regression test. Reported by: jrtc27 Reviewed by: jrtc27 MFC after: 1 week Differential Revision: https://reviews.freebsd.org/D58050
inotify: Ensure that "allocfail" is initialized in inotify_log_one() Fixes: https://cgit.freebsd.org/src/commit/?id=b70997c8c75a ("inotify: Unconditionally generate IN_IGNORED events for files/dirs")
jaildesc_alloc() finishes initializing the file structure only after it
is made visible from the file descriptor table via finit(). In that
window, other threads could try to perform operations on the descriptor
and thus access an incompletely initialized jaildesc.
Defer the finit() call until locks are initialized. While here,
simplify the error path for falloc_caps().
Reported by: Yuxiang Yang, Yizhou Zhao, Ao Wang, Xuewei Feng, Qi Li,
and Ke Xu from Tsinghua University using GLM-5.2 from Z.ai
Reviewed by: jamie
MFC after: 1 week
Sponsored by: The FreeBSD Foundation
Differential Revision: https://reviews.freebsd.org/D58049
dtrace: Improve DOF section size validation The loop which validates each DOF section assumes that the section header is present, so the section size must be at least as large as the header, otherwise a small OOB access is possible. Reviewed by: christos MFC after: 2 weeks Sponsored by: CHERI Research Centre Differential Revision: https://reviews.freebsd.org/D57975
dtrace: Fix DOF section bounds validation We must ensure that each DOF section does not overlap with the DOF header or section table. Otherwise the relocations processed in the second pass over sections can manipulate DOF metadata, leading to OOB writes. Reviewed by: christos MFC after: 2 weeks Sponsored by: CHERI Research Centre Differential Revision: https://reviews.freebsd.org/D57976
dtrace: Improve DOF string table validation The check for a nul terminator implicitly assumes that the section size is positive. Make the assumption explicit. Reviewed by: christos MFC after: 2 weeks Sponsored by: CHERI Research Centre Differential Revision: https://reviews.freebsd.org/D57977
dtrace: Fix DOF section-specific validation The entry size of the probe section is assumed to be at least sizeof(dof_probe_t) by the loop further below. enoff_sec->dofs_entsize was not being validated at all. When multiplying an index by a table entry size, make sure the multiplication can't overflow. Fix an off-by-one when validating the translated probe argument array. Make sure that the probe argument argvs are valid string offsets even if the argument count is zero. Reviewed by: christos MFC after: 2 weeks Sponsored by: CHERI Research Centre Differential Revision: https://reviews.freebsd.org/D57979
Fixes: https://cgit.freebsd.org/src/commit/?id=2ec2ba7e232d ("vfs: Add VFS/syscall support for Solaris style extended attributes") Reported by: Yuxiang Yang, Yizhou Zhao, Ao Wang, Xuewei Feng, Qi Li, and Ke Xu from Tsinghua University using GLM-5.2 from Z.ai Reviewed by: rmacklem, kib MFC after: 1 week Sponsored by: The FreeBSD Foundation Differential Revision: https://reviews.freebsd.org/D58053
A test site determined that, for a Mellanox NIC which can handle M_EXTPG mbufs, an improvement of 5-15% for read rate could be achieved if the read reply was in M_EXTPG mbufs. A patch that tried to determine if the outbound NIC supported M_EXTPG mbufs (IFCAP_MEXTPG) did not pass review. However, it does appear that this can be useful for NFS-over-RDMA. (Which just happen to use NICs that do support M_EXTPG mbufs.) As such, this patch enables them is xp_extpg is set to true, which is never for now, but might be set true for RDMA or when vfs.nfsd.enable_mextpg is set non-zero. (It is 0 by default, so this is never enabled by default at this time.) Tested by: Greg Becker <becker.greg@att.net> MFC after: 2 weeks
Check that the P_WEXIT flag is set. Requested by: markj Reviewed by: markj Tested by: pho Sponsored by: The FreeBSD Foundation MFC after: 1 week Differential revision: https://reviews.freebsd.org/D57124
sys_procdesc: extract procdesc_alloc() Reviewed by: markj Tested by: pho Sponsored by: The FreeBSD Foundation MFC after: 1 week Differential revision: https://reviews.freebsd.org/D57124
sys_procdesc: extract pdtofdflags() Reviewed by: markj Tested by: pho Sponsored by: The FreeBSD Foundation MFC after: 1 week Differential revision: https://reviews.freebsd.org/D57124
sys_procdesc: extract procdesc_destroy() Reviewed byL markj Tested by: pho Sponsored by: The FreeBSD Foundation MFC after: 1 week Differential revision: https://reviews.freebsd.org/D57124
Introduce pd_fpcount that counts the number of file references to the procdesc. Remove the PDF_CLOSED flag, now it is expressed as pd_fpcount == 0. Only send SIGKILL and clear pointers when we are closing the last file referencing procdesc. This should be nop until the next commit. Reviewed by: markj Tested by: pho Sponsored by: The FreeBSD Foundation MFC after: 1 week Differential revision: https://reviews.freebsd.org/D57124
Reviewed by: markj Tested by: pho Sponsored by: The FreeBSD Foundation MFC after: 1 week Differential revision: https://reviews.freebsd.org/D57124
Reviewed by: markj Tested by: pho Sponsored by: The FreeBSD Foundation MFC after: 1 week Differential revision: https://reviews.freebsd.org/D57124
Reviewed by: markj Tested by: pho Sponsored by: The FreeBSD Foundation MFC after: 1 week Differential revision: https://reviews.freebsd.org/D57124
This is backward ABI-compatible, because the only place in kernel that uses the structure, namely the mlx5_ib_cq.c:mlx5_ib_create_cq() function, copies in as much structure members as provided by userspace. Tested by: Wafa Hamzah <wafah@nvidia.com> Sponsored by: Nvidia networking MFC after: 1 month
Import Linux upstream commit 3411f9f01b76bd88aa6e0e013847ab6479cb4f24. rdma_umap_priv_init() takes a reference on the rdma_user_mmap entry for every VMA it maps, but rdma_umap_close() never dropped it. The entry was therefore never freed and lingered in ucontext->mmap_xa, tripping WARN_ON(!xa_empty(&ucontext->mmap_xa)) at context teardown and leaking the firmware UAR on every context close. Reviewed by: kib Tested by: Wafa Hamzah <wafah@nvidia.com> Sponsored by: Nvidia networking MFC after: 1 month
Modern Linuxes don't use ethX for almost 15 years already, see [1] and [2]. The translation logic has always been a source of bugs and PITA. Switch default to not translate (long due!) and schedule removal of the code for FreeBSD 17. [1] https://systemd.io/PREDICTABLE_INTERFACE_NAMES/ [2] https://www.freedesktop.org/software/systemd/man/latest/systemd.net-naming-scheme.html Reviewed by: iwtcex_gmail.com, vvd, melifaro, dchagin Differential Revision: https://reviews.freebsd.org/D57852
sendfile: stop abusing kern_writev() Provide convenient wrapper kern_filewrite() around fo_write(). Switch to use it in vn_sendfile(). This allows to avoid duplicate fget() when we already have the reference to the file, which creates a correctness race with the userspace. Also td_retval[0] clearing hack can be removed. Reviewed by: glebius, markj Sponsored by: The FreeBSD Foundation MFC after: 1 week Differential revision: https://reviews.freebsd.org/D58035
kern_filewrite(): unconditionally calculate cnt, it is used by callers Reported by: dhw, madpilot Tested by: dhw Sponsored by: The FreeBSD Foundation MFC after: 1 week Fixes: https://cgit.freebsd.org/src/commit/?id=dfad790c8cca ("sendfile: stop abusing kern_writev()")
kern_writefile(): fix several regressions sendfile(): for trailers uio, set uio_rw to UIO_WRITE instead of checking it kern_filewrite(): remove unused argument offset kern_writev(): the check should compare cnt against zero, not uio_resid Reported by: markj Fixes: https://cgit.freebsd.org/src/commit/?id=dfad790c8cca ("sendfile: stop abusing kern_writev()") Sponsored by: The FreeBSD Foundation MFC after: 1 week
Reviewed by: markj Tested by: pho Sponsored by: The FreeBSD Foundation MFC after: 1 week Differential revision: https://reviews.freebsd.org/D57163
Reviewed by: markj Sponsored by: The FreeBSD Foundation MFC after: 1 week Differential revision: https://reviews.freebsd.org/D57163
Reviewed by: markj Tested by: pho Sponsored by: The FreeBSD Foundation MFC after: 1 week Differential revision: https://reviews.freebsd.org/D57163
Reviewed by: markj Tested by: pho Sponsored by: The FreeBSD Foundation MFC after: 1 week Differential revision: https://reviews.freebsd.org/D57163
Order them alphabetically. Remove redundand sys/param.h. Reviewed by: markj Sponsored by: The FreeBSD Foundation MFC after: 1 week
linux_sys_futex() does not copyin a timespec for the timeout if the
operation is LINUX_FUTEX_TRYLOCK_PI, presumably because it doesn't make
sense to specify a timeout for a try-lock operation. However, this
means that we pass a userspace timespec pointer to
linux_umtx_abs_timeout_init().
Modify linux_futex_lock_pi() to not initialize the timeout if we're
try-locking.
Reviewed by: kib, dchagin
Reported by: Yuxiang Yang, Yizhou Zhao, Ao Wang, Xuewei Feng, Qi Li,
and Ke Xu from Tsinghua University using GLM-5.2 from Z.ai
MFC after: 1 week
Sponsored by: The FreeBSD Foundation
Differential Revision: https://reviews.freebsd.org/D58061
This is a mask, so the new value should have taken the next bit to avoid breaking a shell script that's interpreted by a binmisc-activated interpreter. Add a brief note that the new value is only used within the ELF activator. Fixes: https://cgit.freebsd.org/src/commit/?id=389c124fecb0 ("imgact_elf.c indicate that interpreter [...]") Reported by: "polyduekes" on discord, madpilot Reviewed by: kib, sjg (both previous version) Differential Revision: https://reviews.freebsd.org/D58063
pm_runtime_resume_and_get is used by new versions of amdgpu, and began use between Linux kernel version 6.12, and 6.14. Reviewed by: dumbbell Sponsored by: The FreeBSD Foundation Differential Revision: https://reviews.freebsd.org/D57463
This catches up with 692b0ef1506ba which added CAPENABLED to clock_nanosleep(). Curiously recent additions of the pdopenpid(2) and pddupfd(2) were done before the cited commit, and that regen did not included the change. Sponsored by: The FreeBSD Foundation
Reviewed by: jfree MFC after: 1 week Sponsored by: The FreeBSD Foundation Differential Revision: https://reviews.freebsd.org/D58084
The taskqueue thread loop tries to avoid entering and exiting net epoch read sections for every task. This reduces the overhead of net epoch integration, but the implementation wasn't bounding the length of the read section, so a busy taskqueue thread could hold an epoch open for an unbounded period. This is easy to achieve with the epair task, for instance. Bound the number of tasks that we'll execute without observing the global epoch, and provide a sysctl to control it. Let the default bound be eight. Reviewed by: glebius MFC after: 2 weeks Differential Revision: https://reviews.freebsd.org/D58031
This function is supposed to wait until all pending callbacks have been executed. This is useful in some contexts where we tear down some context (like a VNET jail and its associated UMA zones) synchronously, and we want to make sure that all pending asynchronous callbacks (which may free objects to said UMA zones) have run first. The implementation schedules a callback on each CPU and waits for them all to run. This assumes that, on a given CPU, callbacks are executed in the order that they are pushed. This assumption depends on the implementation of epoch_call_task() and ck_epoch_poll_deferred(), and it is not true in general. Callbacks are pushed onto a per-CPU stack in LIFO order. ck_epoch_poll_deferred() first pulls out the callbacks from epoch - 2, which are always safe to execute, and in so doing reorders them such that the oldest callback as at the top of the stack, so in this case, epoch_call_task() will execute them in order. However, ck_epoch_poll_deferred() may determine that it is safe to execute callbacks from epoch - 1 (or even from the current epoch if there are no active readers), and in this case it will push those callbacks onto the returned stack. This means that epoch_call_task() will invoke those newer destructors before the older ones, which means that epoch_drain_callbacks() may return early. Fix the correctness problem by simply doing all of this twice: once the first callback is invoked, we know that all of the callbacks that were pending at the time that epoch_drain_callbacks() was called are scheduled to be executed, so when the second callback is executed we know that they must be finished. This is slow, but it is already slow, and the slowness is less noticeable after commit dce56594991. I note that in an ideal world, this function would not exist, and all of the teardown would happen asynchronously, rather than the current mismash of synchronous and asynchronous cleanup. PR: https://bugs.freebsd.org/bugzilla/show_bug.cgi?id=290201 Reviewed by: glebius MFC after: 1 month Differential Revision: https://reviews.freebsd.org/D58030
The former is called by the latter. We return NULL because linuxkpi does not implement ACPI (pseudo?) devices associated to regular devices. The amdgpu DRM driver started to use `ACPI_COMPANION()` in Linux 6.13. Reviewed by: emaste Sponsored by: The FreeBSD Foundation Differential Revision: https://reviews.freebsd.org/D57577
It takes a task state as its last argument. We enforce that this state is `TASK_UNINTERRUPTIBLE` for the time being because other states are not interpreted. Change `usleep_range()` to call `usleep_range_state()` with the state set to `TASK_UNINTERRUPTIBLE`, which is what Linux does too. The amdgpu DRM driver starte to use `usleep_range_state()` in Linux 6.13. Reviewed by: emaste Sponsored by: The FreeBSD Foundation Differential Revision: https://reviews.freebsd.org/D57579
Reviewed by: markj Sponsored by: The FreeBSD Foundation MFC after: 1 week Differential revision: https://reviews.freebsd.org/D58147
Currently no filesystems support it. Reviewed by: mckusick Discussed with: markj Tested by: pho Sponsored by: The FreeBSD Foundation MFC after: 1 week Differential revision: https://reviews.freebsd.org/D57658
On amd64 there was 4 bytes of padding between the 20-byte p_comm and (for LP64) 8-byte p_sysent, so the addition of p_execblock just caused that padding to be eaten up. However, on i386, there was no such padding, and so the addition of p_execblock rippled through to p_emuldata. Fixes: https://cgit.freebsd.org/src/commit/?id=e1a84b7708c2 ("execve_block(): a mechanism for mutual exclusion with execve() on the process")
The note type wakes up when there is something for pdwait(2) to report on the process descriptor. Reviewed by: markj Sponsored by: The FreeBSD Foundation MFC after: 1 week Differential revision: https://reviews.freebsd.org/D58123
We need to wake up the pdwait(2) waiters when procdesc event is reported. Reviewed by: markj Sponsored by: The FreeBSD Foundation MFC after: 1 week Differential revision: https://reviews.freebsd.org/D58172
Convert several callers to use fget_procdesc(). Eliminate procdesc_find() and directly use fget_procdesc() in sys_pdkill(). Previous code structure required to fdrop() procdesc while the process is locked. Reviewed by: markj Sponsored by: The FreeBSD Foundation MFC after: 1 week Differential revision: https://reviews.freebsd.org/D58117
LinuxKPI: add system_percpu_wq In Linux v6.17 system_wq was replaced (renamed to) system_percpu_wq, with the old name still present. We just alias system_percpu_wq to linux_system_short_wq like we do for system_wq to keep both around for the forseeable future. Note: the original system_wq was a per-cpu queue upstream as well based on my understanding but we never implemented it as such. That means we are still lacking a per-cpu implementation for system_percpu_wq but at least we do not change the status-quo of the LinuxKPI implementation with this. Note2: we should add a check somewhere for LINUXKPI_VESION >= 61700 to print a warning if anyone still uses the system_wq to detect any possible sami-native or out-of-tree drivers relying on this and not properly updating. Sponsored by: The FreeBSD Foundation MFC after: 3 days Reviewed by: dumbbell; emaste (comments on previous review) Differential Revision: https://reviews.freebsd.org/D57730
LinuxKPI: fix lkpi_pci_get_device() reference counting on device In case we are passed an "odev" (a device to start the search from), that device would have an extra reference. The best way to illustrate this is to look at for_each_pci_dev(), which will return one device after the other. Upon first return we return a pdev with a reference. That pdev is then passed in as odev on the next call. If we do not clear the reference it will be leaked. Sponsored by: The FreeBSD Foundation MFC after: 3 days Fixes: https://cgit.freebsd.org/src/commit/?id=910cf345d0ee9 ("LinuxKPI: pci: implement ...") Reviewed by: dumbbell, emaste Differential Revision: https://reviews.freebsd.org/D57428
LinuxKPI: pci detach: implement a proper detach (release) path There are two paths in the LinuxKPI PCI code to instantiate a "pdev" (LinuxKPI pci_dev). One is using the FreeBSD bus framework and the pdev will be the softc. This commit starts cleaning up the detach path for just that case to the best possible. So far we did a lot of the work in linux_pci_detach_device(), which is the internal handler of the detach function and little in the (*release) callback (devres cleanup only). The problem with that is, that we tear down resources which later in the devres cleanup are needed. With them not being there anymore we panic, e.g., in lkpi_dma_unmap < lkpi_dmam_free_coherent < lkpi_devres_release_free_list. The solution is to migrate most of the cleanup work into the (*release) callback, which will automatically be called when the device (kobj) reference drops to zero. The only work which should be done immediately is to let the dirver do its cleanup; this has to happen before we try to teardown the resources, but also we do want this to happen when detach is called (the first time). One problem we have with the deferred cleanup of the remaining parts is that we do not know upon calling pci_dev_put() whether this cleared the last reference and triggered the cleanup or not but we cannot return from the detach function with pending resources and dangling pointers, which then may be used. In order to work around this, we clear the (*release) callback function when it is run and check for that in the detach routine. If the (*release) callback was not run, we refuse to detach (force would be needed) as we'd rather keep the device than risk a follow-up panic on leaked resources. Given this should not happen in a well programmed world, I believe it is fine to take that and log it to let the user know. Try to leave a few comments behind to help with understanding in the future. With this we can unload the mt7921 driver (or shutdown the system) without panic. Sponsored by: The FreeBSD Foundation MFC after: 3 days Reviewed by: dumbbell Differential Revision: https://reviews.freebsd.org/D57429
LinuxKPI: 802.11: lkpi_80211_txq_tx_one() only pass sta if added to drv If we are doing a direct (*tx) downcall, only pass sta as meta data if it was added to the driver (via the state machine). This prevents us passing a sta not known to the driver leading to possible follow-up complications/errors. This will usually happen if (a) we are doing software scanning, or (b) if net80211 decides to change the ni from under us and sends a packet with the new ni. Adjust a debug statement before to also have the added_to_drv field in it to ease debugging. Sponsored by: The FreeBSD Foundation MFC after: 3 days
LinuxKPI: pci: fix dma handle type in match function dma_addr_t is a vm_paddr_t which is a uint of some width. Rather than passing pointers of it around pass the value. Comparing the addresses of different storage for the same dma handle (the actual bug here) will not work when passed to the devres match function. Sponsored by: The FreeBSD Foundation Fixes: https://cgit.freebsd.org/src/commit/?id=0a575891211ef ("implement dmam_free_coherent()") MFC after: 3 days Differential Revision: https://reviews.freebsd.org/D58285
LinuxKPI: sg_page() remove superfluous () Sponsored by: The FreeBSD Foundation MFC after: 3 days Reviewed by: emaste Differential Revision: https://reviews.freebsd.org/D58295
LinuxKPI: move clear_page() within the linux/page.h file clear_page() would normally live in asm/page.h but adding the file and fixing the dependencies would be too much for a single line. Move the function to the end of the file with a clear separator and make it clear that it does not operate on a 'struct page' but on a page address by changing the argument name and leaving a comment. The function is currently used by at least mthca(4) as the only in-tree consumer, and drm-kmod ttm_pool.c. No functional changes. Sponsored by: The FreeBSD Foundation MFC after: 3 days Reviewed by: emaste Differential Revision: https://reviews.freebsd.org/D58296
LinuxKPI: prefer struct page [*] over struct vm_page[_t] LinuxKPI is based on Linux 'struct page' which is currently aliased to struct vm_page. Upcoming changes may change that so start using 'struct page *' instead vm_page_t to make future changes transparent. This is a continuation of 9e9c682ff3a1 and should be a NOP. Sponsored by: The FreeBSD Foundation MFC after: 3 days Reviewed by: emaste (no objections) Differential Revision: https://reviews.freebsd.org/D58297
LinuxKPI: page.h: use atop() and ptoa() instead of PAGE_SHIFT With upcoming changes to 'struct page' this will make the lines easier to read by using the predefined macros from param.h. Sponsored by: The FreeBSD Foundation MFC after: 3 days Reviewed by: markj, kib Differential Revision: https://reviews.freebsd.org/D58298
LinuxKPI: page.h: resort lines Two of the "page macros" can be abstracted elsewhere in the upcoming struct page work, so sort them away from the four which are here to stay. No functional change. Sponsored by: The FreeBSD Foundation MFC after: 3 days Reviewed by: emaste Differential Revision: https://reviews.freebsd.org/D58299
LinuxKPI: page pool updates and add to the build Split implementation out from the header files. This "page pool" is the very minimalistic version we need in order to support packets on mt76. We allocate the page pool in order to have the meta data available of which we only make limited use. This implementation does no pooling, it does no page fragments for now, it always hands out a full page and frees it upon return. It is written in a way that it can be in the tree before the 'struct page' work it depends on has landed in order to reduce friction for people who want to try mt7921 (or others later) upfront. We use the same #ifdef as in the struct page work for that reason so one knob will turn everything on or off. Once the struct page work has landed and settled we can start filling this with more complexity. In the unlikely event that in the mean time any other consumer would start showing up they will have to be aware that the current code as-is essentially is a NOP without the 'struct page' work. A WARN_ONCE() will notify them. Sponsored by: The FreeBSD Foundation MFC after: 3 days
LinuxKPI: rcu: add optional condition to list_for_each_entry_rcu() list_for_each_entry_rcu() can take an optional condition. Add the macro argument so code remains compiling but do not do anything with it just yet. Leave comments as list_for_each_entry_rcu() likely should have a different implementation. Sponsored by: The FreeBSD Foundation MFC after: 3 days Reviewed by: dumbbell Differential Revision: https://reviews.freebsd.org/D59292
LinuxKPI: add #include of rculist.h to ethtool.h for Linux conformance Add #include rculist.h to ethtool.h to fullfill expectations of Linux drivers without having to modify them. MFC after: 3 days Reviewed by: dumbbell, emaste Differential Revision: https://reviews.freebsd.org/D58874
LinuxKPI: add typedef for clockid_t Needed by a wireless driver (not really using it) to compile. MFC after: 3 days Reviewed by: dumbbell Differential Revision: https://reviews.freebsd.org/D58877
LinuxKPI: pci: add pci_select_bars() and pci_msix_vec_count() For pci_select_bars() we use the LinuxKPI internal information, while for pci_select_bars() we fall back to the native PCI stack. Needed by a wifi driver. MFC after: 3 days Reviewed by: dumbbell Differential Revision: https://reviews.freebsd.org/D58878
LinuxKPI: fix argument type to lkpi_pci_msi_desc_alloc() lkpi_pci_msi_desc_alloc() takes an unsigned int, not an int. While here make sure the prototype is visibile in interrupt.h as well before use to avoid -Wimplicit-function-declaration errors. Discovered while working on a wireless driver. MFC after: 3 days Reviewed by: dumbbell Differential Revision: https://reviews.freebsd.org/D58879
LinuxKPI: implement dma_{alloc,free}_{noncoherent,attrs}
We use one to implement the other given direction stays unused.
This is needed by an upcoming wifi driver.
MFC after: 3 days
Reviewed by: dumbbell
Differential Revision: https://reviews.freebsd.org/D58880
LinuxKPI: add PCI_IRQ_AFFINITY #define Add the #define for PCI_IRQ_AFFINITY and leave a pr_debug note in pci_alloc_irq_vectors() that it is unimplemented. The flag is needed by an upcoming WiFi driver. MFC after: 3 days Reviewed by: dumbbell Differential Revision: https://reviews.freebsd.org/D58881
LinuxKPI: pci: add reset_{prepare,done} to pci_error_handlers
Needed by an upcoming wireless driver.
MFC after: 3 days
Reviewed by: dumbbell
Differential Revision: https://reviews.freebsd.org/D58882
LinuxKPI: add can_wakeup option and accessor functions We can implement device_set_wakeup_capable() in the !CONFIG_PM_SLEEP case; we do not have the infrastructure in place for the CONFIG_PM_SLEEP case so leave a pr_debug TODO. Needed by an upcoming wireless driver. MFC after: 3 days Reviewed by: dumbbell Differential Revision: https://reviews.freebsd.org/D58883
Checking hlt_cpus_mask is a no-op, and the mask will be removed in the next commit. However, we can use the more recent CPU_ABSENT() macro to check the status. Reviewed by: olce MFC after: 1 week Sponsored by: The FreeBSD Foundation Differential Revision: https://reviews.freebsd.org/D58157
It is a relic, apparently once populated by a machdep.hlt_cpus sysctl. The sysctl was removed, and ULE has never honored this mask. It is now safe to remove. Remove the mask, and its few remaining references in: sched_4bsd(4), hwpmc(4), and hwt(4). Reviewed by: olce, kib MFC after: 1 week Sponsored by: The FreeBSD Foundation Differential Revision: https://reviews.freebsd.org/D58158
The check is always true, especially after the removal of hlt_cpus_mask from sched_4bsd. Reviewed by: olce, kib MFC after: 1 week Sponsored by: The FreeBSD Foundation Differential Revision: https://reviews.freebsd.org/D58159
This makes an effort to clarify and correct the intent of the code,
which is to either:
1. Create one software crypto worker thread for each CPU, to be pinned
later
2. Create the number of threads requested by the kern.geom.eli.threads
tunable
This is as described in geli(8).
If a CPU were somehow* absent, it should be skipped, but not in the
second case when creating a set number of threads.
To achieve this cleanly and correctly:
- split worker creation logic into a helper function
- keep the loops separate
- debug message for absent CPUs is dropped
- add a short explanatory comment
- style, rename local var to 'nthreads'
*Practically, it is impossible today to get a bootable system with a
sparsely populated CPU map. Thus these concerns are hypothetical and
this change should have no functional effect.
Finally, while here, guard the sc->sc_workers list insertion with the
appropriate mutex. The code is safe from races today, but this gives a
better guarantee.
Reviewed by: kib
MFC after: 1 week
Sponsored by: The FreeBSD Foundation
Differential Revision: https://reviews.freebsd.org/D58214
Like the rest of <acpi/video.h>, this function is unimplemented and returns `-ENODEV`. The amdgpu DRM driver started to use it in Linux 6.13. Reviewed by: bz Sponsored by: The FreeBSD Foundation Differential Revision: https://reviews.freebsd.org/D57576
Reviewed by: kib Effort: CHERI upstreaming Sponsored by: DARPA, AFRL Differential Revision: https://reviews.freebsd.org/D58055
exterr: allow exterr to fit pointers on CHERI targets Switch to uint64ptr_t which is a uint64_t on traditional architectures and a uintptr_t on CHERI architectures. This has no ABI impact on non-CHERI kernels. Fix truncation of 64-bit values on 32-bit kernels. Reviewed by: kib Effort: CHERI upstreaming Sponsored by: Innovate UK Differential Revision: https://reviews.freebsd.org/D58056
exterr_set: sync the definition with the header declaration This unbreaks buildkernel with TARGET=armv7 (32-bit arm). More work may be required in order to unbreak `exterr_set` with 32-bit kernels. Fixes: https://cgit.freebsd.org/src/commit/?id=844009378da9 ("exterr: allow exterr to fit pointers on CHERI targets")
kern: fix compilation uintptr64_t -> uint64ptr_t Fixes: https://cgit.freebsd.org/src/commit/?id=5cafd6213f145 (exterr_set: sync the definition with the header declaration)
Remove dependency on sys/proc.h. Reviewed by: imp Sponsored by: Innovate UK Differential Revision: https://reviews.freebsd.org/D58235
m_unshare() had crashed if unmapped mbufs exist in the mbuf chain. This was because memcpy() with mtod() was used without making sure that the mbuf was mapped. Use m_copydata() that cares unmapped mbufs instead. Reviewed by: gallatin Differential Revision: https://reviews.freebsd.org/D58189
We do this already for ET_REL files, but it was missed here. Note that this function operates only on dynamically loaded files, not on preloaded files. Reviewed by: kib MFC after: 2 weeks Sponsored by: The FreeBSD Foundation Differential Revision: https://reviews.freebsd.org/D58245
Sponsored by: The FreeBSD Foundation MFC after: 1 week
Sponsored by: The FreeBSD Foundation MFC after: 1 week
Sponsored by: The FreeBSD Foundation MFC after: 1 week
Sponsored by: The FreeBSD Foundation MFC after: 1 week
Reviewed by: markj Sponsored by: The FreeBSD Foundation MFC after: 1 week Differential revision: https://reviews.freebsd.org/D58247
Since malloc(9) even with M_NOWAIT is forbidden when we hold a spinlock, we can't print detailed lock tree as the operation tries to allocate memory. Fixes: https://cgit.freebsd.org/src/commit/?id=fb4b0c91195195561560bb2fb2c1ba8da81f7ccf
clock_gettime(CLOCK_TAI) can fail, leaving *ovalue uninitialized. Reported by: Hazley Samsudin of GovTech CSG MFC after: 3 days Sponsored by: The FreeBSD Foundation Differential Revision: https://reviews.freebsd.org/D58225
A vm_page's a.queue field records the page queue index for the page queue to which the page belongs. The PGA_ENQUEUED flag indicates whether the page is actually enqueued in that queue's TAILQ. When modifying the a.queue field, you need to hold the page queue lock for the queue corresponding to the old value, unless the old value is PQ_NONE. Suppose a managed page is freed. vm_page_free_prep() calls vm_page_dequeue_deferred(), which checks whether the page belongs to a queue; if so it schedules an asynchronous dequeue operation so that page queue lock acquisitions can be batched if possible. The dequeue operation must be completed before the page's plinks.q fields are reused. So, during page allocation, we call vm_page_dequeue() to finish the dequeue operation. Similarly, since the buddy allocator uses the plinks.q fields for its own internal linkage, vm_freelist_add() calls vm_page_dequeue(). _vm_page_pqstate_commit_dequeue() is the function which actually removes the page from its queue. It sets a.queue = PG_NONE and removes the page from its queue. However, the update to the page's atomic state is relaxed, so on systems with store reordering, it may race with a concurrent enqueue of the page into the buddy queues (probably more likely) or a page queue. Fix this: use a release store to update the page's queue state in _vm_page_pqstate_commit_dequeue(), and make sure that vm_page_dequeue() uses an acquire load when comparing m->a.queue == PQ_NONE. PR: https://bugs.freebsd.org/bugzilla/show_bug.cgi?id=296767 Reported and tested by: pkubaj Reviewed by: alc, kib MFC after: 1 week Differential Revision: https://reviews.freebsd.org/D58261
LinuxKPI: skbuff: improve debugging Deal with SKB_TRACE_FMT optional arguments; while here properly indent. Add KASSERT to __skb_unlink() to catch incorrect skbuffs encountered while debugging a wireless driver (which had other pre-conditions failing). Sponsored by: The FreeBSD Foundation MFC after: 3 days
LinuxKPI: skbuff: add skb_put_zero() Add skb_put_zero() as a simple wrapper around __skb_put_zero(). Sponsored by: The FreeBSD Foundation MFC after: 3 days
LinuxKPI: skbuff: implement napi_build_skb() Implement napi_build_skb() around linuxkpi_build_skb(). Sponsored by: The FreeBSD Foundation MFC after: 3 days
LinuxKPI: skbuff: implement __skb_linearize() skb_linearize() is used by mt7921, mt7925, and in the general mt76 tx dma code. __skb_linearize() is used in the general iwlwifi TX code but given the way we currently create TX skbs in LinuxKPI 802.11 we never hit that case. Sponsored by: The FreeBSD Foundation MFC after: 3 days
LinuxKPI: skbuff: add support for frags in linuxkpi_skb_copy() Sponsored by: The FreeBSD Foundation MFC after: 3 days
LinuxKPI: skbuff: add reference counting to the skb Sponsored by: The FreeBSD Foundation MFC after: 3 days
LinuxKPI: skbuff: add initial page pool support Add an internal flag which is set by skb_mark_for_recycle() and upon "skb_free" then selects whether the skb is freed or returned to the page pool. There will likely be more details to figure out once the LinuxKPI page work is done and we support more of the page pool than the bare minimum. Sponsored by: The FreeBSD Foundation MFC after: 3 days
rtw89(4) would constantly try to start a TX BlockACK session even if no HT or higher was available. The only way to stop this (currently) is to return -EINVAL instead of any other error. Note: we should investigate if/when to call (*set_tid_config)() as that will also offer the ability to forbid BA. Sponsored by: The FreeBSD Foundation Reported by: arved, bnovkov Tested by: bnovkov MFC after: 3 days
vm_phys: Add a sysctl to dump registered fictitious memory ranges I've wanted this a couple of times in the past. Save the memattr in the fictitious memory segment structure so that we can report it from the sysctl handler, and add conversion routines for each platform. Reviewed by: kib MFC after: 2 weeks Differential Revision: https://reviews.freebsd.org/D58283
arm64: Fix the build Fixes: https://cgit.freebsd.org/src/commit/?id=a7e483ee146a ("vm_phys: Add a sysctl to dump registered fictitious memory ranges")
vm: Make sure NULL is defined for vm_memattr_name() Fixes: https://cgit.freebsd.org/src/commit/?id=a7e483ee146a ("vm_phys: Add a sysctl to dump registered fictitious memory ranges")
Reviewed by: markj Sponsored by: The FreeBSD Foundation MFC after: 1 week Differential revision: https://reviews.freebsd.org/D58334
Reviewed by: markj Sponsored by: The FreeBSD Foundation MFC after: 1 week Differential revision: https://reviews.freebsd.org/D58292
procdesc: report NOTE_PDSIGCHLD for traced and stopped process on attach of the knote. It is same as for NOTE_EXIT when attaching to the exiting process. Reviewed by: markj Sponsored by: The FreeBSD Foundation MFC after: 1 week Differential revision: https://reviews.freebsd.org/D58327
kqueue: Fix delivery of unwanted events In both procdesc_kqops_event() and filt_proc(), the event variable can have more than one bit set. This means that: * We cannot compare it directly with NOTE_EXIT; we must binary-and them instead. * We cannot binary-or it with the report mask; we must binary-and it with the request mask first. MFC after: 1 week Fixes: https://cgit.freebsd.org/src/commit/?id=2a5e58c59694 ("procdesc: add NOTE_PDSIGCHLD") Fixes: https://cgit.freebsd.org/src/commit/?id=b328975b9d7c ("procdesc: report NOTE_PDSIGCHLD for traced and stopped process") Reviewed by: kib, markj Differential Revision: https://reviews.freebsd.org/D58395
When transferring a thread with near 100% CPU statistics (but not 100%; up to 57.5/59≈97.46%) to a CPU where the enqueue offset is ahead of at least 2 from the dequeue one, which requires peculiar conditions to happen (transfer triggered by a bind request or cpuset change, or during balancing if a thread or more existed from a brief amount of time on the origin CPU), the transferred thread can get placed after the dequeue offset, effectively making it appear as a high priority one unduly, causing latency increase for other threads. The change here was missed when changing the enqueue and dequeue offsets update mechanism to recover pre-256-queue-runqueue ULE anti-starvation and fairness behavior. That change opened up the possibility that these two offsets are apart by more than one. Reviewed by: markj Discussed with: Minsoo Choo <minsoo@minsoo.io> Fixes: https://cgit.freebsd.org/src/commit/?id=6792f3411f6d ("sched_ule: Recover previous nice and anti-starvation behaviors") MFC after: 2 weeks Sponsored by: The FreeBSD Foundation Differential Revision: https://reviews.freebsd.org/D57829
Different command sets have different encoding for op codes, etc. While one can normally puzzle out which is which, it's better to explicitly tag the command set used. Sponsored by: Netflix
ptrace(2): add PT_GET_CHILDREN Reviewed by: markj Sponsored by: The FreeBSD Foundation MFC after: 1 week Differential revision: https://reviews.freebsd.org/D58315
kern/sys_ptrace: do not skip P2_PTRACEREQ wait for PT_CLEARSTEP/PT_GET_CHILDREN Reported and reviewed by: markj Fixes: https://cgit.freebsd.org/src/commit/?id=d3b7bbee9275 ("ptrace(2): add PT_GET_CHILDREN") Sponsored by: The FreeBSD Foundation MFC after: 1 week Differential revision: https://reviews.freebsd.org/D58364
amd64: FRED support FRED support as defined starting from the SDM rev. 90, requires a new 'events' entry point to receive user and kernel mode exceptions and interrupts notifications from the hardware. A minimal asm trampoline is enough, rest can be implemented in C due to the clean FRED organization of the event reporting. The syscall entry is handled by a microptimized assembly path, directly calling into the amd64_syscall() handler, instead of the generic events entry point. Tested by: emaste Sponsored by: The FreeBSD Foundation MFC after: 1 week Differential revision: https://reviews.freebsd.org/D55829
amd64: Fix an off-by-one in the fred_ipi_handlers definition Fixes: https://cgit.freebsd.org/src/commit/?id=6e93f5e4d693 ("amd64: FRED support") Reviewed by: kib Differential Revision: https://reviews.freebsd.org/D58378
amd64: Remove a prototype for an unimplemented function Fixes: https://cgit.freebsd.org/src/commit/?id=6e93f5e4d693 ("amd64: FRED support") Reviewed by: kib Differential Revision: https://reviews.freebsd.org/D58379
Return the covered vnode instead. Tested by: pho Sponsored by: The FreeBSD Foundation MFC after: 1 week Differential revision: https://reviews.freebsd.org/D58191
We introduced (PRI_MAX_TIMESHARE - PRI_MIN_TIMESHARE) as part of
ESTCPULIM() in commit eebc148f25c3 ("sched_4bsd: ESTCPULIM(): Allow any
value in the timeshare range") in order to use more than a fixed number
(40) of all the available priority levels in the timeshare range (136
before the 256-queue runqueue work, 224 now) to take into account the
number of ticks a thread has run ('ts_estcpu').
In the computation of a new thread's priority (resetpriority()), in
addition to the "ticks running" contribution, the final priority also
includes a "nice" value contribution. The final value is clamped into
the [PRI_MIN_TIMESHARE; PRI_MAX_TRIMESHARE] range.
Problem is that the new "ticks running" contribution now can lead to
a computed priority value that exceeds PRI_MAX_TRIMESHARE, and is thus
finally clamped to PRI_MAX_TIMESHARE, which becomes an alias for all
out-of-bound values. In particular, this can conflate CPU-hungry
threads. With at least two of them competing on the same CPU, with an
increase of 'ts_estcpu' of ~64 per second (stathz being 127) and the
minimal decay of 4/5 (load average 2 or more), both threads will easily
reach the current clamping of 224 (+ PRI_MIN_TIMESHARE), and be
considered indifferently by the scheduler.
Fix this problem by ensuring that the maximum contribution of
'ts_estcpu' (via ESTCPULIM()) cannot exceed the timeshare range of
priorities when the nice contribution is added to it, so the nice
contribution continues to have an effect on CPU-bound threads.
Introduction of the nice term in ESTCPULIM() (then NICE_WEIGHT *
PRIO_MAX) has been done in commit bdf423572ee3 ("Scheduler fixes
equivalent to the ones logged in the following NetBSD commit...") and
does not appear to have made any real sense even then.
Fixes: https://cgit.freebsd.org/src/commit/?id=bdf423572ee3 ("Scheduler fixes equivalent to the ones logged in the following NetBSD commit...")
Fixes: https://cgit.freebsd.org/src/commit/?id=eebc148f25c3 ("sched_4bsd: ESTCPULIM(): Allow any value in the timeshare range")
MFC after: 2 weeks
Sponsored by: The FreeBSD Foundation
Differential Revision: https://reviews.freebsd.org/D57826
The INVERSE_ESTCPU_WEIGHT scaling had been introduced by commit
b698380f33ef ("Quick fix for scaling of statclock ticks in the SMP
case. ...") to leave more discrimination room for multiple CPUs possibly
adding their ticks to the same 'struct ksegrp' (but also slightly
changing how CPU hogs are penalized).
Then, commit 8460a577a4b4 ("Make KSE a kernel option, ...") introduced
the current thread-based code, where tick accounting is only done on the
current thread, which renders this trick obsolete on !KSE.
Finally, when KSE was removed, the trick became generally obsolete.
The trick is actually even harmful because it changes the intended
behavior of priorizing more the CPUs that use the less ticks (and so,
impairs boosting "interactive" processes).
Remove it now. Clamping of 'ts_estcpu' and its relation to the
load-average-based decay may be re-examined later.
Fixes: https://cgit.freebsd.org/src/commit/?id=8460a577a4b4 ("Make KSE a kernel option, ...")
MFC after: 2 weeks
Sponsored by: The FreeBSD Foundation
Differential Revision: https://reviews.freebsd.org/D57827
In an upcoming change whose purpose is to stop having 4BSD always
allocate MAXCPU runqueues, wasting space on most machines, 'struct
td_sched' will store the CPU ID to which a thread is bound/pinned
instead of a pointer to the corresponding runqueue. As a consequence,
existing functions manipulating a thread's runqueue will need to point
to the inferred runqueue through a local variable. The name 'runq' is
the ideal one for these local variables, but before this change it
designated the global runqueue, also causing unnecessary ambiguity.
Thus, rename the global runqueue to the more explicit 'runq_global'.
Arguably, this should have been performed as part of commit e17c57b14ba9
("- Implement cpu pinning and binding. (...)").
No functional change (intended).
[olce: Massaged the commit message. Tested with source builds.]
Suggested by: olce
Reviewed by: olce
Tested by: olce
MFC after: 2 weeks
Differential Revision: https://reviews.freebsd.org/D58065
4BSD has been allocating an array of MAXCPU runqueues, runq_pcpu[],
instead of one runqueue per actually present CPU. On amd64, MAXCPU is
1024 and 'struct runq' is 4128 bytes, causing runq_pcpu[] to take more
than 4 MiB of memory. On the vast majority of current systems, which
have at most 32 cores with SMT, this is a waste of memory.
Besides providing per-CPU runqueues, runq_pcpu[] has also been used to
determine the CPU ID of a given thread's associated runqueue through
pointer arithmetic.
Since per-CPU structures are only allocated for present CPUs, in order
to save space, move the runqueues to per-CPU fields and, for each thread
('struct ts_sched'), replace its runqueue pointer by the CPU ID of the
runqueue it is in (new 'ts_rqcpu' field). Set the thread's CPU ID to
the special NOCPU value when it is running on the global runqueue.
Drop the SKE_RUNQ_PCPU() macro as it is now simply equivalent to
'ts_rqcpu != NOCPU'. Introduce the TS_RUNQ_PTR() macro to get a pointer
to the thread's runqueue, which must be passed to runq_add() and
runq_remove().
[olce: Massaged the commit message. Fixed an inverted KASSERT().
Tested with source builds.]
Reviewed by: olce
Tested by: olce
MFC after: 2 weeks
Differential Revision: https://reviews.freebsd.org/D58000
We would lock the downcalls during normal operation but not during vap (vif) creation as there was no need for locking. Add the missing locking there as drivers seem to always expect it (by assertion) and cannot distinguish between state. Add the assertions to the downcalls as we need both of them locked and both of them can sleep. PR: https://bugs.freebsd.org/bugzilla/show_bug.cgi?id=296185 ("rtwhttps://bugs.freebsd.org/bugzilla/show_bug.cgi?id=89(https://bugs.freebsd.org/bugzilla/show_bug.cgi?id=4) freezes the system with INVARIANTS kernel") Debugged by: Artem Bunichev (temcbun gmail.com) Sponsored by: The FreeBSD Foundation MFC after: 3 days
Allow userland, in particular test cases for EXTERROR conversions, to detect at run time whether extended errors include the descriptive message strings, which depends on the EXTERR_STRINGS kernel option and cannot be probed in any other way. Reviewed by: kib MFC after: 1 week Assisted-by: Claude Code (Fable 5) Differential Revision: https://reviews.freebsd.org/D58321
On a test system with 1024 cores the size of exec map exceeds 4GB, and all of the operands in the size calculation are 32-bit integers. Tested by: Jim Huang Chen <jim.chen.1827@gmail.com> MFC after: 1 week Sponsored by: AMD (hardware)
Also be more protective in getsid(). Reported by: arrowd Reviewed by: markj Sponsored by: The FreeBSD Foundation MFC after: 1 week Differrential revision: https://reviews.freebsd.org/D58393
Fixes: https://cgit.freebsd.org/src/commit/?id=963629923308 ("kthread_add(): do not allow to attach the thread to a dead or dying process") Reviewed by: kib MFC after: 1 week Sponsored by: The FreeBSD Foundation Differential Revision: https://reviews.freebsd.org/D58433
Reported by: Maik Muench of Secfault Security Reviewed by: kib Fixes: https://cgit.freebsd.org/src/commit/?id=1ad21a652182 ("kern: add pddupfd(2)") Sponsored by: The FreeBSD Foundation Differential Revision: https://reviews.freebsd.org/D58403
The FD_RESOLVE_BENEATH flag is supposed to be sticky. It's set when you receive an fd from a different jail and preserved by openat(<dfd>) etc.. However, if you send the fd to yourself, the flag is stripped since SCM_RIGHTS message don't preserve file descriptor flags. Fix this by preserving those flags and checking for UF_RESOLVE_BENEATH in restrict_rights(). Fixes: https://cgit.freebsd.org/src/commit/?id=350ba9672a7f ("unix: Set O_RESOLVE_BENEATH on fds transferred between jails") Reviewed by: kib MFC after: 1 week Sponsored by: The FreeBSD Foundation Differential Revision: https://reviews.freebsd.org/D58317
As far as I can see, it is impossible for procdesc_exit() to observe pd->pd_fpcount == 0: if procdesc_close() decrements that counter to zero, then it will clean up the procdesc structure too, and this is atomic with respect to the proctree lock. No functional change intended. Reviewed by: kib MFC after: 1 week Sponsored by: The FreeBSD Foundation Differential Revision: https://reviews.freebsd.org/D58396
The scan marker was originally stack-allocated. In commit 1c0f9af5b5224, it became heap-allocated since the marker is visible to other threads and a scanning thread's stack may be swapped out. Now that kernel stacks can no longer be swapped out, we can avoid these heap allocations. Reviewed by: kib MFC after: 1 week Sponsored by: The FreeBSD Foundation Differential Revision: https://reviews.freebsd.org/D58402
uma: Factor out the implementations of uma_zfree_{arg,smr}()
The two function both free an item to a UMA zone, but uma_zfree_arg()
does so in such as way as to ensure that the item will be the first one
returned by a subsequent allocation, while uma_zfree_smr() must defer
reuse of the item and therefore never frees to the per-CPU alloc bucket.
When KASAN is enabled, we actually want uma_zfree_arg() to behave like
uma_zfree_smr(): to improve the reliability of use-after-free detection,
reuse of the newly freed item should be deferred for some time.
Refactor a bit to make it easier to improve KASAN along these lines:
introduce two helper functions, cache_free_item() and cache_free_smr(),
which handle most of the work of interacting with the per-CPU caches.
A subsequent commit will let uma_zfree_arg() use cache_free_smr() when
KASAN is enabled.
No functional change intended.
Reviewed by: rlibby
MFC after: 1 month
Sponsored by: The FreeBSD Foundation
Differential Revision: https://reviews.freebsd.org/D58268
uma: Make an effort to defer reuse of items when KASAN is enabled When KASAN is configured, make uma_zfree_arg() free items to the per-CPU free bucket, rather than to the alloc bucket. This means that the item won't be recycled immediately the next time a thread goes to allocate an item from that zone on the same CPU. In other words, the item will stay in a quarantine state longer, which helps make KASAN's use-after-free detection more reliable. Reviewed by: rlibby MFC after: 1 month Sponsored by: The FreeBSD Foundation Differential Revision: https://reviews.freebsd.org/D58269
uma: Avoid allocating from free buckets when KASAN is enabled When uma_zalloc_arg() hits an empty alloc bucket in the per-CPU cache, it tries swapping the alloc and free buckets in the hope that the free bucket has some items available. If not, it has to lock the zone. Disable this behaviour when KASAN is configured in order to further defer reuse of freed items. This forces a free item to go to the per-domain full bucket cache before it becomes accessible to the allocator. Reviewed by: rlibby MFC after: 1 month Sponsored by: The FreeBSD Foundation Differential Revision: https://reviews.freebsd.org/D58270
uma: Enqueue full buckets in FIFO order when KASAN is configured We want to defer reuse of free objects, and this is a trivial way to promote that. Suggested by: rlibby Reviewed by: rlibby, alc MFC after: 1 month Sponsored by: The FreeBSD Foundation Differential Revision: https://reviews.freebsd.org/D58312
This makes it easier to grep for the error message to better understand the call stack when loading firmware modules fails. Fix a cosmetic-only style(9) bug while here in the same function related to another logging message. MFC after: 1 week Differential Revision: https://reviews.freebsd.org/D58380
If something goes very badly (e.g. forcibly removing a medium while the OS tries to start it), this could end up in params.blksize being 0 (and params.disksize 1). Avoid an integer divide fault, panicking the kernel, by bailing out before. MFC after: 3 days
Add a NOTE_REAP event for EVFILTER_PROC which provides a notification when the process is reaped. MFC after: 1 week Sponsored by: Klara, Inc. Sponsored by: NetApp, Inc. Reviewed by: kib, markj Differential Revision: https://reviews.freebsd.org/D58313
Reviewed by: markj Sponsored by: The FreeBSD Foundation MFC after: 1 week Differential revision: https://reviews.freebsd.org/D58463
This matches the documented prototype and avoids spurious -Wincompatible-pointer-types-discards-qualifiers warnings when passing a constant pathname. Sponsored by: AFRL, DARPA
PR: https://bugs.freebsd.org/bugzilla/show_bug.cgi?id=297062 Tested by: Jordan Gordeev <jgopensource@proton.me> Reviewed by: markj Sponsored by: The FreeBSD Foundation MFC after: 1 week Differential revision: https://reviews.freebsd.org/D58472
p_reapsubtree lives in the p_startcopy/p_endcopy block of struct proc, which is copied during fork without any synchronization. However, the field is not stable except when the proctree lock is held, and indeed may change if p1's reaper exits or explicitly releases its reaper status. This state change can race with fork() and leave the child with an incorrect p_reapsubtree field. Close the race: explicitly copy the field under the proctree lock during fork. Reported by: syzkaller Reviewed by: kib MFC after: 2 weeks Sponsored by: The FreeBSD Foundation Differential Revision: https://reviews.freebsd.org/D58482
kqueue: Add a helper macro for sleeping on in-flux knotes Other in-flux operations are implemented by this set of macros, so we should do the same for sleeping. No functional change intended. Reviewed by: kib MFC after: 1 week Sponsored by: The FreeBSD Foundation Differential Revision: https://reviews.freebsd.org/D58443
kqueue: Associate marker knotes with a queue Otherwise the assertion in KQ_FLUX_SLEEP_WMESG may fail. kqueue_fork_copy() already handles this. Fixes: https://cgit.freebsd.org/src/commit/?id=1f4b0ea4f3eb ("kqueue: Add a helper macro for sleeping on in-flux knotes") Reported by: syzkaller Reported by: kbowling Reviewed by: kib Sponsored by: The FreeBSD Foundation Differential Revision: https://reviews.freebsd.org/D58516
Otherwise an assertion in umtx_thread_alloc() (TAILQ_EMPTY(&uq->uq_pi_contested)) is violated. This use of TDB_EXIT is hacky, but I cannot see another way to check for an exiting thread without adding some more overhead to kern_thr_exit(). Fixes: https://cgit.freebsd.org/src/commit/?id=2a339d9e3dc1 Reported by: Maik Muench of Secfault Security Reviewed by: kib MFC after: 1 week Sponsored by: The FreeBSD Foundation Differential Revision: https://reviews.freebsd.org/D58447
MFC after: 1 week Sponsored by: Klara, Inc. Sponsored by: NetApp, Inc.
proc_realparent(): assert that an orphaned child has real parent != parent Reviewed by: markj Sponsored by: The FreeBSD Foundation MFC after: 1 week Differential revision: https://reviews.freebsd.org/D58504
proc_realparent(): do not mark the child as orphan when reparenting to p_opptr pid Reported and reviewed by: markj Fixes: https://cgit.freebsd.org/src/commit/?id=8cef3c9b768a ("proc_realparent(): assert that an orphaned child has real parent != parent") Sponsored by: The FreeBSD Foundation MFC after: 1 week Differential revision: https://reviews.freebsd.org/D58566
to avoid using uninitialized value in the KASSERT() statement on the first iteration. Also, do the assert under the proctree_lock, which is not critical but satisfies the invariants. Noted and reviewed by: markj Sponsored by: The FreeBSD Foundation MFC after: 1 week Differential revision: https://reviews.freebsd.org/D58505
We reach ndaasync with the CAM device lock held, so we must pass M_NOWAIT to disk_* rather than M_WAITOK. Reviewed by: imp Fixes: https://cgit.freebsd.org/src/commit/?id=628d7a3270b6 ("nda: AC_GETDEV_CHANGED calls media chanaged for sectorsize change") MFC after: 1 week Sponsored by: Amazon Differential Revision: https://reviews.freebsd.org/D58230
Grow the module collection to answer, each with a single command, the first questions asked when diagnosing a sick system: why is my application stalling, where is the kernel fighting over locks, what file could it not find, why is this process getting EPERM, what killed my process, will that fatal signal actually leave a core behind, what was my process stuck on, where is my kernel memory going, who is creating or entering jails, is the network slow because TCP is resending, how long did my thread wait to run, and is the disk itself slow. Every module keeps to the house style: invocation-name overloading through hard links, predicate-only D with inline lookup tables (no if-statements), and stable providers only (syscall, proc, sched, io, dtmalloc, and the lockstat, vfs, priv, and mib SDT providers), so the modules remain drop-in compatible with older releases (the one documented exception is noted below). No kernel changes: new and extended profiles under cddl/usr.sbin/dwatch/libexec plus one libdtrace inline table (priv.d). slow (slow-fsync, slow-open, slow-read, slow-syscall, slow-write, or any slow-NAME by new link) records syscall entry timestamps in thread-local storage and prints, at return, any call whose latency meets a threshold (DWATCH_SLOW_MS, default 100), naming the syscall, the elapsed time to the microsecond, and any errno returned. The bare profile watches a curated set of filesystem-related calls expected to be fast; slow-syscall watches everything; unrecognized invocation names fall through to syscall::NAME:return with the matching entry probe derived mechanically from the return probe list. lock (lock-adaptive, lock-block, lock-lockmgr, lock-rw, lock-spin, lock-sx, lock-thread) rides the dtrace_lockstat(4) block and spin probes, printing the held-off thread (free from the standard event tag), the holdoff duration, the lock class, the lo_name of the lock through a single cast of arg0 to struct lock_object (the first member of every kernel lock), and reader/writer intent on the probes that report it. Holdoffs shorter than DWATCH_LOCK_MS (default 1; 0 shows everything) are suppressed. namei (namei-enoent, namei-entry, namei-failure) records the pathname at vfs:namei:lookup:entry and reports it with the result at return. Unlike the vop_lookup profile, which reconstructs paths from the name cache one component at a time, this sees the whole path exactly as the process requested it. namei-enoent hunts file-not-found storms -- the single most common use of truss(1) -- without stopping the victim. priv (priv-err, priv-ok) watches priv_check(9) verdicts, naming the exact privilege denied -- something no syscall tracer can see, because by the time EPERM surfaces the priv(9) value is gone. The number is decoded by priv_string[], a new libdtrace inline table in the errno.d and signal.d tradition, mechanically generated from sys/priv.h (247 entries) and installed to /usr/lib/dtrace where dtrace(1) auto-loads it; on older releases it is a drop-in file like the module itself. coredump (coredump-top) watches for delivery of signals whose default action produces a core, per the SIGPROP_CORE entries of the sigproptbl in kern_sig.c, and renders a verdict the same way and in the same order the kernel will decide it: ignored or caught per the target's struct sigacts, then the coredump() gauntlet of kern.coredump, kern.sugid_coredump vs P_SUGID, procctl(2) PROC_TRACE_CTL, and RLIMIT_CORE -- the sysctl knobs read live through kernel globals. Where a coredump-worthy signal will produce no core, the verdict says precisely which policy ate it. coredump-top maintains a cumulative catalog of coredump-worthy signals by process and signal, refreshed every 3 seconds in the style of systop; combine the event profile with `-O cmd' to capture state as each event occurs. hang (hang-top) pairs sched:::sleep with sched:::wakeup through a tid-keyed timestamp array and prints, as each thread wakes, any sleep that meets a threshold (DWATCH_HANG_MS, default 1000), naming the sleeper in the details and the waker in the standard event tag. This is the blocking the slow module structurally cannot see: a syscall that never returns never reports its latency, while hang reports the moment the wait ends, with the full duration. hang-top maintains a cumulative catalog of long sleeps by process (count and maximum) in the style of coredump-top. jail (jail-attach, jail-get, jail-remove, jail-set) watches the jail management plane -- jail(2), jail_set(2), jail_get(2), jail_attach(2), and jail_remove(2) -- naming the operation, the jail id (taken from the entry argument for attach/remove, from the return value for the others), and any errno. Complements the dwatch `-j jail' filter, which scopes any profile to processes inside one jail; this watches who manipulates jails, from any jail or none. dtmalloc (dtmalloc-top, or any dtmalloc-NAME by new link) rides the dtmalloc provider (one malloc and one free probe per malloc(9) type). The event profile prints allocations and frees meeting a size threshold (DWATCH_MALLOC_MIN, default 65536) -- who is allocating huge kernel buffers. dtmalloc-top maintains a running catalog of net bytes and outstanding allocation balance by type, sorted by net bytes so leak suspects rise: a type that climbs without bound while the system is in steady state is the suspect. The catalog reflects activity since the watch began, and is honest about caches holding what they allocate. mib (tcp-retransmit, or any mib-NAME by new link) rides the per-counter mib SDT probes of the network stack. The tcp-retransmit profile curates the counters that signal send-path congestion or loss -- data packet retransmissions, unnecessary retransmissions, retransmit timer expirations, and connections dropped by retransmit exhaustion -- decoded through an inline description table, answering "is this network slow because TCP is resending?" as events with process context rather than netstat(1) deltas. NB: the mib probes exist only in kernels built with options KDTRACE_MIB_SDT (default in -CURRENT via std.debug); the module documents this and dtrace(1) refuses the script elsewhere, making the dependency self-announcing. Four existing modules gain personalities. proc grows proc-signal-fatal, filtering signal-send to signals whose default disposition terminates the receiver, most-notably including kernel-generated SIGSEGV/SIGBUS/SIGILL/SIGFPE that no kill(2) watcher will ever see. errno now reads its invocation name: errno-NAME shows only syscalls returning that errno, where NAME is a symbolic name from errno.d or a number; links are installed for errno-EACCES, errno-ECAPMODE, errno-ENOENT, errno-ENOTCAPABLE, and errno-EPERM (the latter pairs covering capsicum(4) capability-mode violations), and any other errno needs only a new link. sched grows sched-latency, recording a timestamp at sched:::enqueue keyed by tid and printing at sched:::on-cpu any run-queue wait meeting a threshold (DWATCH_SCHED_MS, default 10) -- the literal measurement of scheduler delay on a system with idle CPU that still feels sluggish. io grows io-slow, pairing io:::start with io:::done through a bio-keyed timestamp array and printing any request that meets a threshold (DWATCH_IO_MS, default 100), naming the device, command, size, and elapsed time; watched against zvols and a pool's leaf vdevs this brackets where in a ZFS stack the time is going, without touching unstable providers. Document all of the above plus the DWATCH_HANG_MS, DWATCH_IO_MS, DWATCH_LOCK_MS, DWATCH_MALLOC_MIN, DWATCH_SCHED_MS, and DWATCH_SLOW_MS knobs in dwatch(1). All 46 new invocation names were exercised through `dwatch -d' with a profile-path sandbox emulating the installed hard links: every one sources cleanly and emits the intended D -- probe selection per alias, entry/return and sleep/wakeup pairing through thread-local and global associative arrays, threshold and mask predicates picking up their knobs, aggregation clauses and printa column layout in the -top profiles, multi-line predicate rendering, and `-t' correctly displacing each module's default test were verified by inspection of the generated scripts. Invocations untouched by this pass generate D identical to their previous output. Modules pass sh -n, fit 80 columns, and dwatch.1 passes mandoc -Tlint with no new warnings. A validation harness performing a `dwatch -e' compile per profile against the live kernel globs every staged profile for runs wherever the dtrace device is present. Reviewed by: markj Differential Revision: https://reviews.freebsd.org/D58093
Reviewed by: jah, markj Sponsored by: The FreeBSD Foundation MFC after: 1 week Differential revision: https://reviews.freebsd.org/D58477
There are probably more places which could benefit from allowing to interrupt vfs_busy() calls at syscalls top level. Requested by: Peter Eriksson <pen@lysator.liu.se> Reviewed by: jah, markj Sponsored by: The FreeBSD Foundation MFC after: 1 week Differential revision: https://reviews.freebsd.org/D58477
Otherwise we'll print an error but carry on regardless, presumably destined to walk off the end of the mapping. Reported by: thebugfixers@pm.me MFC after: 1 week
knotes with a non-trivial f_copy implementation may be activated before kqueue_fork_copy_knote() is finished. In particular, it may be enqueued at the time that kqueue_fork_copy_knote() calls knote_enqueue(). Guard against this. Add a test case which triggers the race. Fix several other problems with the replication of knote state: - Make sure only the KN_ACTIVE and KN_DISABLED status flags are inherited, the rest should not be copied. - Ignore marker knotes. - Ignore knotes for kqueues. They cannot be safely copied into the child without more work, as kqueues are inherently local to a process; on fork, we need to ensure that such knotes are patched to reference the new kqueue, not the original. - Try to keep knote state stable by holding the kqueue and knlist locks while copying. Approved by: so Security: FreeBSD-SA-26:50.kqueue Security: CVE-2026-58083 Reviewed by: kib Reported by: Hazley Samsudin of GovTech CSG Sponsored by: The FreeBSD Foundation Differential Revision: https://reviews.freebsd.org/D58223
Approved by: so Security: FreeBSD-SA-26:52.if_wg Security: CVE-2026-58085 Reviewed by: markj Sponsored by: Chelsio Communications
Commit 4be491e1b9b3 ("jail: Optionally allow audit session state to
be configured in a jail") removed the #if 0 around the audit cases
in prison_priv_check() and added the PR_ALLOW_SETAUDIT check under
them. This unintentionally captured the preceding case PRIV_KTRACE,
which used to fall through the disabled block into the unconditional
return (0) of the credential cases: since then, jailed root only has
ktrace privileges (tracing processes with changed credentials, see
ktrcanset()) when the unrelated allow.setaudit knob is enabled, and
conversely gains them when that audit knob is turned on.
Give PRIV_KTRACE back its own unconditional return (0), matching its
comment and the pre-4be491e1b9b3 behaviour.
Approved by: so
Security: FreeBSD-SA-26:53.ktrace
Security: CVE-2026-58086
Fixes: https://cgit.freebsd.org/src/commit/?id=4be491e1b9b3 ("jail: Optionally allow audit session state to be configured in a jail")
Reviewed by: markj
Assisted-by: Claude Code (Fable 5)
These commands take a snapshot of the size of a semaphore set, then drop
the lock and malloc an appropriately sized array before reacquiring the
lock. A comment explains why this is (probably) safe. Unfortunately,
it's wrong; it is indeed possible for a malicious userspace to create
and destroy 2^{15} sets in the window where the lock is dropped. This
race can lead to out-of-bounds reads and writes, and that can be
exploited to elevate privileges.
Replace the assertions with runtime checks.
Approved by: so
Security: FreeBSD-SA-26:54.sysvsem
Security: CVE-2026-58087
Reported by: Maik Muench of Secfault Security
Reviewed by: kib
Sponsored by: The FreeBSD Foundation
Differential Revision: https://reviews.freebsd.org/D58421
In an ELF coredump, each dumped vm_map_entry is represented by a segment. __elfN(coredump) first computes the number of segments by looping over the vm_map entries (in each_dumpable_segment()), then allocates a buffer to hold the ELF header and program headers, then loops over the entries again to populate the program headers. each_dumpable_segment() holds the vm_map read lock, but that lock is dropped between the two calls. If the map is shared with another process, via rfork(), then the map can change. cb_put_phdr() did not account for this, and so could write out of bounds. Add a check to prevent this; simply do not write out excess segments. Approved by: so Security: FreeBSD-SA-26:55.elf Security: CVE-2026-58088 Reported by: Maik Muench of Secfault Security Reviewed by: kib, emaste Sponsored by: The FreeBSD Foundation Differential Revision: https://reviews.freebsd.org/D58416
Reviewed by: kib Obtained from: Linux commit d41861942fc55c14b6280d9568a0d0112037f065 Sponsored by: Chelsio Communications Differential Revision: https://reviews.freebsd.org/D57952
Otherwise ktls_mbuf_crypto_state() will reject mbufs created by _mb_unmapped_to_ext(), which arises when transmitting packets through an interface that doesn't support unmapped mbufs, and the loopback interface in particular. PR: https://bugs.freebsd.org/bugzilla/show_bug.cgi?id=296498 Fixes: https://cgit.freebsd.org/src/commit/?id=3444414cb463 ("ktls: Don't attempt to modify non-anonymous mbufs on the receive path") Reviewed by: gallatin, jhb MFC after: 1 week Sponsored by: The FreeBSD Foundation Differential Revision: https://reviews.freebsd.org/D57557
Reported by: Nick Price Tested by: pho Reviewed by: jah, markj Sponsored by: The FreeBSD Foundation MFC after: 1 week Differential revision: https://reviews.freebsd.org/D58506
Commit f2202ab5abda did not account for KTLS mbufs. m_unshare() tries to linearize the original mbuf chain and creates a writable copy of it, converting unmapped mbufs. Both of them are unsafe for KTLS mbufs. It is better to return NULL if the mbuf chain contains a KTLS mbuf. Reported by: jhb Reviewed by: jhb Differential Revision: https://reviews.freebsd.org/D58466
Without this, KASAN has the deficiency that inter-object overflows are not detected most of the time[*] when keg_layout() is able to perfectly pack a slab. Try to overcome this by adjusting the allocation size to include a redzone following the object. With this change, we automatically get a redzone following each item, so any overflow into the redzone will trigger a panic. Most of UMA doesn't need to know about this: at slab allocation time, the whole slab is poisoned, and then kasan_mark_item_valid() will unpoison only the buffer that is available to the consumer. Note that in most zones, most objects will follow another object's redzone, so there is some protection against underflow as well. It might be worthwhile to provide a stronger guarantee here. Add an assertion to item_ctor() that the returned item is properly aligned. I couldn't see any pre-existing checks which verify this. Reviewed by: rlibby MFC after: 2 weeks Sponsored by: The FreeBSD Foundation Differential Revision: https://reviews.freebsd.org/D58271
malloc: Refactor redzone and sanitizer handling malloc_large() duplicates redzone and KASAN handling that is also present in malloc() and malloc_domainset(). Refactor the implementations to reduce this a bit. Also normalize KMSAN map handling: make malloc() and malloc_domainset() consistent, and do not update the KMSAN shadow map, as we can rely on UMA and kmem_malloc() to handle that. Reviewed by: rlibby MFC after: 3 weeks Sponsored by: The FreeBSD Foundation Differential Revision: https://reviews.freebsd.org/D58272
uma: Fix KMSAN integration with malloc zones In commit 459aa032e872 I dropped kmsan_mark() calls from malloc() on the basis that UMA and kmem_malloc() would handle updates of the KMSAN shadow map. However, I missed that UMA explicitly does not handle this. Modify UMA to only omit origin map updates for malloc zones. Fixes: https://cgit.freebsd.org/src/commit/?id=459aa032e872 ("malloc: Refactor redzone and sanitizer handling") Reviewed by: rlibby Differential Revision: https://reviews.freebsd.org/D58574
We should of course pass the provided domainset rather than copying what plain malloc() does. Fixes: https://cgit.freebsd.org/src/commit/?id=89deca0a3361 ("malloc: make malloc_large closer to standalone") Reviewed by: rlibby MFC after: 3 weeks Sponsored by: The FreeBSD Foundation Differential Revision: https://reviews.freebsd.org/D58316
kern_proc_kqueues_out() reported into an intermediate sbuf and copied the result into the caller's. A process that had leaked 468k kqueue descriptors wired 757 MB of M_SBUF while dumping core, over roughly 9M reallocations, then copied the whole thing again. Reviewed by: adrian, markj Differential Revision: https://reviews.freebsd.org/D58536 PR: https://bugs.freebsd.org/bugzilla/show_bug.cgi?id=296835 MFC after: 1 week
All sorts of places in the ELF loading code assume that filesz <= memsz, so check that explicitly up front. Reported by: Jane Smith <thebugfixers@pm.me> (via D57785) Reviewed by: jrtc27, kib Differential Revision: https://reviews.freebsd.org/D58542
This just invokes xa_insert similar to other xa_*_irq wrappers. Reviewed by: bz Sponsored by: Chelsio Communications Differential Revision: https://reviews.freebsd.org/D58576
If boot_mute is set the system appears to hang during the mountroot prompt. Temporarily unmute the console so the prompt is visible. Reviewed by: kib MFC after: 1 week Differential Revision: https://reviews.freebsd.org/D58549
Reviewed by: markj Tested by: pho Sponsored by: The FreeBSD Foundation MFC after: 1 week Differential revision: https://reviews.freebsd.org/D58407
Instead of accessing the struct proc and gathering data from it, memoize the data needed for pdwait() on exited process in struct procdesc, at the time of process termination. This allows unlimited number of calls to pdwait(2) on procdesc for terminated process. Change the locking requirements for pd_flags to proctree_lock. This does not modify the pre-patch locking regime, but the change requires it. Reviewed by: markj Tested by: pho Sponsored by: The FreeBSD Foundation MFC after: 1 week Differential revision: https://reviews.freebsd.org/D58407
Add the p_zombieref bitmask into struct proc, which enumerates all legitimate waiters on the process exit status. Among them are parent for PZOMBIEREF_PARENT, and the holder of the process descriptor for PZOMBIEREF_PROCDESC, if the process was created by pdfork(). Require all zombie refs to be cleared to reap zombie. This prevents stealing the exit status from the parent by pdwait()ing on a procdesc obtained by pdopenpid(), or by waitpid() by debugger from the real parent. Reviewed by: markj Tested by: pho Sponsored by: The FreeBSD Foundation MFC after: 1 week Differential revision: https://reviews.freebsd.org/D58264
Reimplement atomic_{set,clear}_16 using atomic_set_32.
Remove emulation of these operations from vm_page.c.
Reviewed by: alc, kib
MFC after: 2 weeks
Differential Revision: https://reviews.freebsd.org/D58580
sbuf reserves a byte of its buffer for the terminator, so the sbuf created with maxlen held one byte less than the sizing pass had computed. The last record overflowed it, sbuf_bcat() failed, and the error == 0 guard skipped the copy into the caller's sbuf, so the note has been emitted at full size but zero filled since 5e7c43ff02dc. Fixes: https://cgit.freebsd.org/src/commit/?id=5e7c43ff02dc Reviewed by: adrian, markj Differential Revision: https://reviews.freebsd.org/D58583 MFC after: 1 week
kern_proc_kqueues_out() sized its intermediate sbuf from the preceding sizing pass, so dumping core for a process with many knotes wired a buffer as large as the entire report. Shrank the intermediate to one page and added a drain that copied into the caller's sbuf up to maxlen, stopping the walk once it was reached. Truncation stayed byte exact. A dump of 384k knotes peaked at 20 KB of M_SBUF instead of 445 MB. Reviewed by: adrian, markj Differential Revision: https://reviews.freebsd.org/D58584 MFC after: 1 week
The "add missing GIDs" loop uses rdma_find_gid_by_port() to test whether a GID already exists, but forgets to drop the reference it returns. So every rescan that finds an existing GID leaks one, which pins the entry and prevents its slot from ever being freed on delete. Just release the reference once the GID is found, like the "remove stale GIDs" loop already does. Reported by: Wafa Hamzah <wafah@nvidia.com> Reviewed by: kib, jhb Sponsored by: Nvidia networking Fixes: https://cgit.freebsd.org/src/commit/?id=6a75471dbcf0 ("OFED: Various changes from Linux 4.19") Differential revision: https://reviews.freebsd.org/D58511
When cleaning up stale GIDs the scan stopped as soon as rdma_get_gid_attr() failed. But that can also happen for empty entries in the middle of the table, so a single gap left everything after it behind and the GID entries could eventually run out. Now the whole table is scanned and the empty slots are simply skipped. Reviewed by: kib, jhb Sponsored by: Nvidia networking Fixes: https://cgit.freebsd.org/src/commit/?id=6a75471dbcf0 ("OFED: Various changes from Linux 4.19") Differential revision: https://reviews.freebsd.org/D58510
The static, global index into an array of strings is simple, cheap, and works for the base kernel, but is unworkable for (potentially third-party) kernel modules or for arbitrary userspace code. Swipe a few of the top bits of category to indicate a source with all-zeros being the current model (EXTERR_CAT_SRC_KERN_STATIC). Add two additional sources EXTERR_CAT_SRC_KERN_DYNAMIC and EXTERR_CAT_SRC_USER with stub implementations. Reviewed by: kib Sponsored by: Innovate UK Differential Revision: https://reviews.freebsd.org/D58236
Make it possible to define categories without compiling their paths into libc (important for third-party modules). The EXTERR_CATEGORY_DYNAMIC macro can be defined to a string describing the compilation unit (generally the path relative to src/sys) which takes the place of EXTERR_CATEGORY. These strings are assembled in linker sets with category numbers assigned at system startup or module load time. The strings can be retrieved from the kern.exterr.categories.<category> sysctl. Reviewed by: kib Sponsored by: Innovate UK Differential Revision: https://reviews.freebsd.org/D58237
AF_MAX was always intended to be one more than the greatest allocated value. Jeff broke this in 2013. Unfortunately, a bunch of people then decided to adapt to the mistake instead of correcting it. Fixes: https://cgit.freebsd.org/src/commit/?id=863c7e45628d (" - Reserve a special AF for SDP. The one we were incorrectly using before was taken by another AF.") MFC after: 3 days Sponsored by: Klara, Inc. Sponsored by: NetApp, Inc. Reviewed by: kevans, glebius Differential Revision: https://reviews.freebsd.org/D58597
Approved by: kib Pull Request: https://github.com/freebsd/freebsd-src/pull/2349
For pdwait(2) and pddupfd(2), the returned error is kept EINVAL. PR: https://bugs.freebsd.org/bugzilla/show_bug.cgi?id=297293 Reviewed by: lwhsu, markj Sponsored by: The FreeBSD Foundation MFC after: 1 week Differential revision: https://reviews.freebsd.org/D58666
Drivers which remap PF queues need a stop/mutate/restart transaction only when the interface has live queues. Permit their IOV initialization callback while the interface is administratively down and leave it down afterward. This restores the standard boot-time iovctl.conf workflow for igb and lets other opt-in drivers configure VFs before netif brings the PF up. MFC after: 1 week
Also add vnode locking wrappers for lockcanrecurse(9) and lockdisablerecurse(9). Reviewed by: jah Sponsored by: The FreeBSD Foundation MFC after: 1 week Differential revision: https://reviews.freebsd.org/D58567
For some complex nullfs mount configurations, it is possible to get the covered vnode lock for the mount shared with some inside-mount vnode lock. Then at unmount time, vflush() would recurse on the covered vnode lock when reclaiming the vnode. Work around it, by temprorarily allowing recursion on the covered vnode lock. Disable recursion after the unmount if it was not enabled before. PR: https://bugs.freebsd.org/bugzilla/show_bug.cgi?id=297174 Reviewed by: jah Sponsored by: The FreeBSD Foundation MFC after: 1 week Differential revision: https://reviews.freebsd.org/D58567
Fixes: https://cgit.freebsd.org/src/commit/?id=ed85203fb7a0 ("vmm: Deduplicate VM and vCPU state management code") Reviewed by: markj MFC after: 1 week Differential Revision: https://reviews.freebsd.org/D58697
It matches our IEEE80211_ELEMID_HTINFO, that is already in use. Found with: clang -Werror=assign-enum
Killing a knote releases its file reference, and releasing the last one runs the close path inline. panic: _mtx_lock_sleep: recursed on non-recursive mutex ttymtx Revoking a controlling tty during exit reaches this whenever a knote is still registered on it. Released the knlist lock around the drop and restart the walk. The knote stays valid while the lock is released. MFC: 1 week Reviewed by: kib Differential Revision: https://reviews.freebsd.org/D58681
iflib counts resets initiated by its transmit watchdog in 69c3e0de01c1. Export the counter in the per-device iflib sysctl tree so every driver provides the diagnostic without a driver callback or duplicate storage. A watchdog reset does not establish how many packets failed. It can recover a hardware stall involving several queued packets or a missed completion involving no packet loss. Stop adding one output error per watchdog event in em(4), igb(4), and igc(4). Remove the redundant driver counters and move the diagnostic to dev.<driver>.<unit>.iflib.tx_watchdog_events. MFC after: 1 month Relnotes: yes
pci: Add SR-IOV status reporting Add a generic packed-nvlist status query to each /dev/iov/<PF> control device. Report the live VF Enable state, configured and total VF counts, and one record for each configured VF. Each VF record contains its PF-local index, computed PCI location, newbus attachment state, attached driver, and ppt binding. Construct records for hardware VFs whose newbus child is absent so attachment failures remain visible. Version the extensible schema in sys/iov.h. Use fixed-width request fields so the ioctl command and layout are identical for 32-bit callers. Serialize the topology snapshot with Giant, then pack and copy it after releasing Giant.
pci_iov: Use native types for status ioctl IOV_CONFIG and IOV_GET_SCHEMA expose native pointers and size_t lengths, and pci_iov has no compat32 ioctl translation. Using fixed-width fields for IOV_GET_STATUS alone does not make the interface usable by 32-bit binaries on a 64-bit kernel. It instead complicates otherwise ordinary pointer and length handling. Use void * and size_t like the existing ioctls. This also makes the %zu diagnostic in iovctl correct on ILP32 and removes the unneeded PTRIN conversion. Fixes: https://cgit.freebsd.org/src/commit/?id=6f8b3be1fbd6 ("pci: Add SR-IOV status reporting")
This change adds wrappers for the new fine-grained TLB invalidation instructions and extends the capability detection logic to include the Svinval extension, which is mandatory in the RVA23S64 profile. Event: BSDCan 2026 Differential Revision: https://reviews.freebsd.org/D57623 Reviewed by: mhorne, markj
unix: Fix a missing initialization in uipc_sosend_stream_or_seqpacket() This could be triggered by an in-kernel sender, of which I can't find any examples. Fixes: https://cgit.freebsd.org/src/commit/?id=d15792780760 ("unix: new implementation of unix/stream & unix/seqpacket") Reviewed by: glebius MFC after: 1 week Sponsored by: The FreeBSD Foundation Differential Revision: https://reviews.freebsd.org/D58673
unix: Simplify uipc_detach() uipc_close() handles detaching a unix socket from the vnode to which it's bound, if any, so doing the same in uipc_detach() is redundant. Moreover, it's conceptually wrong that uipc_detach() might need to handle this: detach happens when there are no remaining references to the socket, and that should include the vnode's reference, even though it's not explicitly counted. No functional change intended. Reviewed by: John Ericson <inquire@JohnEricson.me>, glebius MFC after: 1 week Sponsored by: The FreeBSD Foundation Differential Revision: https://reviews.freebsd.org/D58675
unix: Fix some bugs in the SOCK_STREAM receive path The main problem is with the handling of errors from unp_externalize(). It turns out that this was quite broken, and unfortunately it's easy to trigger such errors (e.g., by setting a low per-process fd limit with setrlimit()). In non-peek mode, uipc_soreceive_stream_or_seqpacket() cuts a bunch of mbufs from the head of the socket buffer, to be consumed by userspace. When unp_externalize() returns an error, we splice the removed mbuf chain back onto the head of the socket buffer. This is expensive, but that's ok since such errors are rare. The problem is that this cutting is not correctly implemented: it does not clear the "next" pointer for the last mbuf in the chain, so it still points to the first mbuf still resident in the socket buffer. This means that mc_init_m() creates a chain that still includes the rest of the socket buffer, so splicing the chain back into the socket buffer does not work properly. Fix this: fully detach the control chain from the socket buffer so that we can safely use mc_init_m(). Then, incrementally add data mbufs, taking care to handle "part". Fix some related bugs while here: - Don't swallow the error if unp_externalize() fails and there's nothing left in the socket buffer (i.e., control->m_next == NULL). - Roll back changes to the partially read mbuf. Reviewed by: glebius MFC after: 1 week Sponsored by: The FreeBSD Foundation Differential Revision: https://reviews.freebsd.org/D58695
unix: split unp_connectat() in two Factor the second half — connecting to an already-resolved peer PCB — out into a new `unp_connect_peer()`, leaving `unp_connectat()` with the connection state machine and pathname resolution. No functional change. The helper's contract: the caller guarantees stability of the peer PCB (vnode lock plus `unp_vp_mtxpool` lock for peers found via `VOP_UNP_CONNECT()`), has set `UNP_CONNECTING` on the connecting socket, and clears it again on error; the helper clears it on success. This prepares for connecting to a peer named by something other than a pathname. Signed-off-by: John Ericson <John.Ericson@Obsidian.Systems> Reviewed by: markj MFC after: 2 weeks Differential Revision: https://reviews.freebsd.org/D58404
unix: factor unp_sun_path() out of bind and connect Extract the AF_UNIX validation plus sun_path/length lookup shared by `uipc_bindat()`, `unp_connect()`, and `unp_connectat()` into a helper that hands back the path pointer and its length. Each caller keeps its own empty-path policy and, where needed, its own copy of the path. Signed-off-by: John Ericson <John.Ericson@Obsidian.Systems> Assisted-by: Claude Code (Claude Opus 4.8 and Fable 5) Reviewed by: markj MFC after: 2 weeks Differential Revision: https://reviews.freebsd.org/D58459
unix: pin the pathname peer by reference across the connect In the pathname path of `unp_connectat()`, take a reference on the peer socket under the per-vnode `unp_vp_mtxpool` lock, drop that lock, and `vput()` the vnode *before* calling `unp_connect_peer()`, rather than holding the vnode lock across the connect. `unp_connect_peer()` already accepts "a reference on the peer socket" as a stability guarantee (it is exactly what the descriptor path relies on), so this is behaviour-preserving. The payoff is that no vnode lock is held across the connect, which removes the delicate `MPASS(!(return_locked && connreq))` "vput() must not sleep while the peer is locked" invariant on the datagram fast path. That reference then has to be released, and for the reasons described in the code, this can only safely happen *after* the PCB is unlocked. The boolean flag is replaced with a nullable out pointer to return the reference to the caller so that it can carry out this responsibility. No functional change intended. Signed-off-by: John Ericson <John.Ericson@Obsidian.Systems> Assisted-by: Claude Code (Claude Opus 4.8 and Fable 5) Reviewed by: markj MFC after: 2 weeks Differential Revision: https://reviews.freebsd.org/D58460
unix: factor unp_vnode_peer() out of unp_connectat() Move the "resolve a locked vnode to the referenced peer socket it names" block into a helper. Pure code motion: the caller now calls `unp_vnode_peer()` and keeps the `vput()`/connect/`sorele()` sequence. No functional change intended. Note: This refactor isn't really necessary as `unp_vnode_peer()` will only be called once throughout this entire patch series. I am just including it out of my personal preferences for decomposing tasks into smaller functions --- we can skip this patch if the reviewers don't like this. Signed-off-by: John Ericson <John.Ericson@Obsidian.Systems> Assisted-by: Claude Code (Claude Opus 4.8 and Fable 5) Reviewed by: markj MFC after: 2 weeks Differential Revision: https://reviews.freebsd.org/D58461
unix: factor unp_connectat_peer() out of unp_connectat() Move the "resolve a connectat(2) target to a referenced peer socket" half of `unp_connectat()` -- the `namei()` lookup and `unp_vnode_peer()` call -- into a helper, leaving `unp_connectat()` with the connection state machine plus a single `unp_connect_peer()`. This is where the next change grows the ways a peer can be named; keeping it a helper up front keeps that change focused on the new resolution logic. No functional change intended. Signed-off-by: John Ericson <John.Ericson@Obsidian.Systems> Assisted-by: Claude Code (Claude Opus 4.8 and Fable 5) Reviewed by: markj MFC after: 2 weeks Differential Revision: https://reviews.freebsd.org/D58462
unix: allow connectat(2) to name the peer socket by descriptor Accept an empty `sun_path` when `fd` is not `AT_FDCWD`: the descriptor then names the peer unix socket directly, instead of being the starting directory for a pathname lookup. The held file reference keeps the peer PCB stable, playing the role `unp_vp_mtxpool` plays in the pathname path. The descriptor must carry `CAP_CONNECTAT` and refer to an `AF_UNIX` socket (`EPROTOTYPE` otherwise, `ENOTSOCK` for non-sockets). As with a pathname, a stream/seqpacket peer must be listening. No filesystem permission or MAC vnode check applies on this path: possession of the descriptor is the authorization, as with descriptor passing. Note this makes it possible to connect a datagram socket to an unbound peer, which no pathname could previously name. `connect(2)` and the implicit-connect send path pass `AT_FDCWD` and still reject an empty path with `EINVAL`. The `unp_sun_path()` call is hoisted out of `unp_connectat()` because the early exit conditions for the two system calls (`connect(2)` and `connectat(2)`) are slightly different. Additionally, support `/dev/fd/<N>`. In a world with `connectat(2)`, this is largely overkill, but this also allows me to add support for direct peer connections with plain `connect(2)`. I think that is a wise choice because this will allow me to propose this functionality for Linux too without a new system call (saving that conversation for later). Ultimately, I want to see multiple operating systems support this to foster broader userland adoption, which should benefit everyone including FreeBSD --- it's nicer if more 3rd party in addition to 1st party software uses the new kernel functionality. Therefore, I hope this additional feature is also acceptable. Signed-off-by: John Ericson <John.Ericson@Obsidian.Systems> Assisted-by: Claude Code (Claude Opus 4.8 and Fable 5) Reviewed by: markj MFC after: 2 months Differential Revision: https://reviews.freebsd.org/D58405
unix: only treat an empty sun_path as a peer descriptor for connectat(2)
connect(2) passes AT_FDCWD to unp_connectat(), so the empty-path
descriptor branch added in 6563dcb6b1f5 turned any sockaddr whose
sun_path begins with a NUL byte into getsock(AT_FDCWD), failing with
EBADF where the pathname lookup historically failed with ENOENT.
Linux abstract namespace names are exactly that: the linuxulator
passes them through with the leading NUL intact, and libxcb tries the
abstract socket first, falling back to the pathname socket only on
ENOENT or ECONNREFUSED. The EBADF made every Linux X11 client fail
at startup with "Missing X server or $DISPLAY".
Restrict the descriptor interpretation to fd != AT_FDCWD, matching
the contract stated in 6563dcb6b1f5's commit message ("Accept an
empty sun_path when fd is not AT_FDCWD"): connect(2) again reaches
the pathname lookup and fails with ENOENT as it always did.
Add a regression test: a NUL-leading, nonzero-length sun_path through
connect(2) or connectat(2) with AT_FDCWD must fail the pathname
lookup with ENOENT, not EBADF.
Fixes: https://cgit.freebsd.org/src/commit/?id=6563dcb6b1f5 ("unix: allow connectat(2) to name the peer socket by descriptor")
Reviewed by: John Ericson <John.Ericson@Obsidian.Systems>, markj
Differential Revision: https://reviews.freebsd.org/D58792
`uipc_listen()` refused a socket that had not been bound, with `EDESTADDRREQ`. That made sense while a pathname was the only way to name a peer: an unbound listener could never be reached, so allowing it would only have created sockets nothing could connect to. Now that `connectat(2)` can name a peer socket by descriptor, an unbound listener *is* reachable, and the restriction only stands in the way. It also left stream sockets oddly stricter than datagram ones, which could already reach an unbound peer. Dropping the check additionally permits `bind(2)` after `listen(2)`: `uipc_bindat()` already allows this, as it only rejects re-binding a socket that has a name. That ordering closes a window listeners otherwise have to leave open. Today the socket file must exist before the socket may listen, so a client connecting in between is refused; binding afterwards publishes the name only once the socket is ready to accept. `unix_seqpacket_test:listen_unbound` asserted the old behaviour, and is inverted accordingly. Signed-off-by: John Ericson <John.Ericson@Obsidian.Systems> Assisted-by: Claude Code (Claude Opus 4.8 and Fable 5) Reviewed by: glebius, markj MFC after: 2 months Differential Revision: https://reviews.freebsd.org/D58683
taskqueue KPI require wakeup() to be called for each completed task. With everything else there heavily optimized over the years, even when doing nothing this wakeup()'s lock/unlock is significant. Since no external taskqueue consumer can depend on the tq_mutex, we can move the wakeup() out of it. It creates some complications for internal waiters, but those should be much more rare, and can be handled with separate locked wakeups on demand. My tests of taskqueue-intensive ZFS RAIDZ writes on 64-core system show performance improvement from this change ~4%, while same time reducing CPU usage by several percent due to lower lock contention, confirmed by CPU profiler.
Reviewed by: markj Differential Revision: https://reviews.freebsd.org/D58711
PR: https://bugs.freebsd.org/bugzilla/show_bug.cgi?id=296348 Reviewed by: markj Sponsored by: The FreeBSD Foundation MFC after: 1 week Differential revision: https://reviews.freebsd.org/D58785
In some private discussion it was pointed out that vm_object_split()'s pattern of dropping the source object lock looks dangerous in that the initial assumption that OBJ_ONEMAPPING is set may become false. In practice I believe that the map lock holds this flag stable, but let's assert that. Reviewed by: alc, kib MFC after: 1 week Differential Revision: https://reviews.freebsd.org/D58766
SYSINIT: add explicit SI_ORDER_LAST Working on cleansing use of (SI_SUB_FOO + 1) construct through the kernel I found a repeating pattern. Often a developer adds a module that depends on certain subsystem to be fully instantiated and they want to put their module SYSINIT right at the end of the SI_SUB_FOO. Such module usually expects that nothing else within this subsystem shall depend on the module. The problem with SI_ORDER_ANY which practically was "the last" until this change is that it is used very widely and people treat it literally as "any", well, because this is what the name says. This lead to many parts that could have dependencies later to be added as SI_ORDER_ANY. So, our developer with the new subsystem that depends on SI_SUB_FOO has three options: 1) Use SI_ORDER_ANY, but grep around ther kernel for other SI_SUB_FOO entries to make sure that no dependencies are set to SI_ORDER_ANY. And in case they are, shift them up and recheck if dependencies of those dependencies are met. 2) Take next subsystem in sysinit list. However, the next one can be SI_SUB_BAR, that is completely irrelevant from SI_SUB_FOO, and our developer doesn't want to put his module's SYSINIT into SI_SUB_BAR, cause it is ugly. 3) Use the (SI_SUB_FOO + 1) construct that violates -Werror=assign-enum. The SI_ORDER_LAST solves this hard choice. If you know that nothing is going to depend on your module within SI_SUB_FOO, but you depend on SI_SUB_FOO, just use SI_ORDER_LAST. Reviewed by: markj, emaste Differential Revision: https://reviews.freebsd.org/D58709
SYSINIT: add SI_SUB_FIRST This allows to initialize mp_maxid, mp_ncpus and register APICs at the most early stage, guaranteeing that those values will already be available at SI_SUB_TUNABLES. Reviewed by: markj Differential Revision: https://reviews.freebsd.org/D58712
SYSINIT: add SI_SUB_NUMA This allows to parse ACPI tables and initialize VM domains before SI_SUB_VM w/o a hack. Reviewed by: markj Differential Revision: https://reviews.freebsd.org/D58713
rtnl_if_flags_to_linux() translated the usual IFF_* bits but dropped FreeBSD's IFF_LOWER_UP (IFF_NETLINK_1). Chromium's AddressTrackerLinux only treats a link as online when ifi_flags has UP|LOWER_UP|RUNNING; with LOWER_UP missing, online_links stays empty, ConnectionType is CONNECTION_NONE, and navigator.onLine is false even though TCP/HTTPS work. Linux Chromium under the Linuxulator (e.g. www/linux-brave) then shows a spurious Offline UI; sites that ignore navigator.onLine do not. Native www/chromium is on a different notifier path. Map IFF_LOWER_UP to Linux's IFF_LOWER_UP (1<<16) and define LINUX_IFF_LOWER_UP alongside the existing LINUX_IFF_* constants. PR: https://bugs.freebsd.org/bugzilla/show_bug.cgi?id=297424 Reviewed by: pouria, adrian (previous revision) Differential Revision: https://reviews.freebsd.org/D58774
video: fix v4l2_buffer size assert on non-i386 32-bit ports Split the #else branch into an explicit __i386__ case (68) and a generic ILP32-with-64-bit-time_t case (80) covering arm and powerpc. Fixes: https://cgit.freebsd.org/src/commit/?id=0343ab8a6afa Reviewed by: adrian Differential Revision: https://reviews.freebsd.org/D58790
video: add V4L2 compat symbols for ffmpeg/opencv Adds v4l2_std_id, struct v4l2_standard/v4l2_plane, VIDIOC_G_STD/S_STD/ ENUMSTD, V4L2_STD_NTSC*, the MPLANE capability flag, multiplanar types (VIDEO_MAX_PLANES, v4l2_plane_pix_format, v4l2_pix_format_mplane, V4L2_TYPE_IS_MULTIPLANAR), V4L2_PIX_FMT_JPEG/YUV411P/SN9C10X, and the MPEG control class (V4L2_CID_MPEG_BASE, V4L2_CID_MPEG_VIDEO_B_FRAMES). video(4) capture devices are digital-only and never expose these, but ffmpeg's libavdevice/v4l2.c and opencv's cap_v4l.cpp both reference them unconditionally. Fixes their build against sys/videoio.h. PR: https://bugs.freebsd.org/bugzilla/show_bug.cgi?id=297454 Reviewed by: manu, adrian Differential Revision: https://reviews.freebsd.org/D58793
video: bump __FreeBSD_version for video(4) refactor Reviewed by: manu, adrian Differential Revision: https://reviews.freebsd.org/D58798
Empty mchains cannot be copied with simple assignment. I think this bug is mostly harmless: if mcnext is empty, then it won't be accessed again before it is reinitialized in the next loop iteration. So the bug only trips an assertion in INVARIANTS kernels and won't be visible otherwise. Add a regression test which triggers this corner case. Reported by: Jan Bramkamp Fixes: https://cgit.freebsd.org/src/commit/?id=d15792780760 ("unix: new implementation of unix/stream & unix/seqpacket") Reviewed by: glebius MFC after: 1 week Sponsored by: The FreeBSD Foundation Differential Revision: https://reviews.freebsd.org/D58791
Fixes: https://cgit.freebsd.org/src/commit/?id=35164034e390 ("arm64/vmm: Make remaining registers use hypctx_*_sys_reg") Reviewed by: markj MFC after: 1 week Differential Revision: https://reviews.freebsd.org/D58690
When the terminal cdev is closed due to revoke, ttydev_close() destroys t_inpoll and t_outpoll selinfos. Since corresponding knotes reference files pointing to the same tty cdev, it fdrop()s them. But then the VOP_CLOSE() call would recurse into the ttydev_close() for the same tty. More, because the devfs vnode is already doomed, each close call gets the FREVOKE flag set. As result, the kernel is recursing as deep into the ttydev_close() as there are opened files referencing the same tty, which have the knotes installed. Basically, the recursion level is controlled by userspace. Prevent it by marking the tty that is handled by ttydev_close(), with the TF_INDEVCLOSE flag. Do nothing in ttydev_close() when the flag is already set, avoiding recursion. Fixes: https://cgit.freebsd.org/src/commit/?id=acd5638e268a ("tty: delete knotes when TTY is revoked") Reviewed by: markj Sponsored by: The FreeBSD Foundation MFC after: 1 week Differential revision: https://reviews.freebsd.org/D58706
This should provide much higher resistence against struct thread layout changes for out-of-tree modules depending on linuxkpi. Reviewed by: bz Sponsored by: The FreeBSD Foundation MFC after: 1 week Differential revision: https://reviews.freebsd.org/D58733
Since vmspace_iop()/proc_readmem() might return -1 on error from vmspace_rwmem(), account for this and stop reading but return already accumulated data if any, instead of returning an error. PR: https://bugs.freebsd.org/bugzilla/show_bug.cgi?id=297512 Reported and tested by: Stéphane D'Alu <sdalu@sdalu.com> Reviewed by: markj Fixes: https://cgit.freebsd.org/src/commit/?id=e1b0d051bbf7 ("proc: Allow to make proc_rwmem() operate on a consistent address space") Sponsored by: The FreeBSD Foundation MFC after: 1 week Differential revision: https://reviews.freebsd.org/D58838
In particular, if there were any bytes moved, and then vm_fault() faulted, do not return an error, but report the short io instead. PR: https://bugs.freebsd.org/bugzilla/show_bug.cgi?id=297512 Reviewed by: markj Tested by: Stéphane D'Alu <sdalu@sdalu.com> Sponsored by: The FreeBSD Foundation MFC after: 1 week Differential revision: https://reviews.freebsd.org/D58838
PR: https://bugs.freebsd.org/bugzilla/show_bug.cgi?id=297516 Reported by: asomers Fixes: https://cgit.freebsd.org/src/commit/?id=dfad790c8cca ("sendfile: stop abusing kern_writev()") Sponsored by: The FreeBSD Foundation MFC after: 3 days
Thanks to Vinicius Ferrao <versatushpc.com.br>, there is now a module that implements the server side of RDMA for the FreeBSD NFS server. At least for now, it will be maintained as an "unofficial port" for FreeBSD, since it was built with generative AI and FreeBSD is working on a policy related to these submissions. This patch puts the "glue" needed by Vinicius's nfsrdma.ko module in the system. This "glue" was written by me without the use of AI. The "unofficial port" of nfsrdma.ko will be advertised on freebsd-current@ as soon as it is available. (Vinicius's work was sponsored by VersatupHPC.) Since newnfs_numnfsd is now declared extern in nfs.h, the extern declaration can be removed from assorted files. I'll do that as a separate commit. Suggested by: Vinicius Ferrao <versatushpc.com.br> MFC after: 1 month
Pre-attach sysctls contain pointers into the iflib context. Any later registration failure that frees the context must first remove that sysctl tree. Failures after a successful IFDI_ATTACH_PRE also did not consistently call IFDI_DETACH or free the private taskqueue. In particular, routing a taskqueue creation failure through the context cleanup could free the driver softc while resources allocated by attach_pre remained live. Track successful interrupt and queue setup and use one common unwind path. Invoke IFDI_DETACH with IFNET_WLOCK dropped and release only resources whose setup completed. Leave a failed IFDI_ATTACH_PRE to unwind its own partial state, as required by the existing driver contract. A failed post-attach can follow driver registration of an SR-IOV schema. Remove that registration before detaching the interface and driver, matching normal deregistration, so a failed attach cannot leave a stale /dev/iov node or make the next attach report EBUSY. A successful attach_pre can now be followed by detach before driver queue allocation. Make the remaining queue-backed interrupt cleanup paths tolerate absent queue arrays. Mark a failed registration as detaching before draining the entire private taskqueue. Drivers can register configuration tasks there, and taskqueue_drain_all() does not wait for work queued during its drain. Make every current non-admin callback reject detaching contexts so late work cannot touch driver state. Drain tasks and call ether_ifdetach() with neither the ifnet nor context lock held. A callback already running may need either lock, while ether_ifdetach() acquires ifnet_detach_sx. Reacquire IFNET_WLOCK before the context lock to preserve the established lock order. The shared automatic core-offset allocator also lacked acquisition state. Late registration failures leaked its reference, while normal detach could decrement a reference belonging to another device when a configured offset or allocation failure meant that this context never acquired one. Record acquisition explicitly and release only references held. MFC after: 2 weeks Reviewed by: gallatin Sponsored by: BBOX.io Differential Revision: https://reviews.freebsd.org/D58721
static, not static inline, so TUs that don't call it warn. Reviewed by: adrian Differential Revision: https://reviews.freebsd.org/D58823
static inline, and powerpc's own must_bounce() never calls it. Reviewed by: adrian Differential Revision: https://reviews.freebsd.org/D58825
linuxkpi: Include <linux/notifier.h> and <linux/device.h> from <linux/pm_qos.h> The i915 DRM driver started to depend on the `bool` type implicitly imported through these headers in Linux 6.13. The previous fix committed in 67df313015906d84d90df8e37795885e81cf8da5 did not reproduce the same includes as Linux. This may have led to other missing implicit definitions later. To prevent another missing include in the future, we already include <linux/plist.h> from <linux/pm_qos.h> and add it as a dummy header. This will be easier to add a proper implementation in the future once we actually need it. Reviewed by: bz, emaste MFC after: 3 days Sponsored by: The FreeBSD Foundation Differential Revision: https://reviews.freebsd.org/D57570
linuxkpi: Define `__GFP_HIGH` in <linux/gfp.h> The DRM drivers TTM memory manager started to use it in Linux 6.15. Reviewed by: emaste Sponsored by: The FreeBSD Foundation Differential Revision: https://reviews.freebsd.org/D58760
linuxkpi: Add `kunit_fail_current_test()` This is part of some unit testing framework. The DRM drivers generic code started to use it in Linux 6.15. Sponsored by: The FreeBSD Foundation Differential Revision: https://reviews.freebsd.org/D58249
linuxkpi: Define `PCI_CLASS_BRIDGE_HOST` The i915 DRM driver started to use it in Linux 6.15. Sponsored by: The FreeBSD Foundation Differential Revision: reviews.freebsd.org/D58250
linuxkpi: Fix return type of `kobject_uevent_env()` The function returns an int on Linux. Let's return 0 (success). Sponsored by: The FreeBSD Foundation Differential Revision: reviews.freebsd.org/D58252
linuxkpi: Create empty <linux/sprintf.h> The DRM drivers generic code started to include this header in Linux 6.15, though nothing from it is used apparently. Sponsored by: The FreeBSD Foundation Differential Revision: https://reviews.freebsd.org/D58253
linuxkpi: Include <linux/page-flags.h> from <linux/mm.h> This reproduces the same namespace pollution as Linux. Sponsored by: The FreeBSD Foundation Differential Revision: https://reviews.freebsd.org/D58761
linuxkpi: Alias `sd` field of `struct kobject`
On Linux, a `struct kernfs_node *` pointer, representing a directory in
sysfs, is kept in the `sd` struct field.
We don't have that on FreeBSD because we use sysctls instead. Let's
alias the sysctl OID pointer to `sd` using an union.
This pointer is checked by the DRM drivers using:
if (var->kobj.sd) {
...
}
The amdgpu DRM driver started to use these checks in Linux 6.13.
Reviewed by: bz
Sponsored by: The FreeBSD Foundation
Differential Revision: https://reviews.freebsd.org/D57586
linuxkpi: Add `kthread_run_worker()` On Linux, `kthread_run_worker()` differs from `kthread_create_worker()` by waking up the task after creating it with `kthread_create_worker()`. On FreeBSD, we already execute it and wait for it, so probably no need to do anything further. Therefore, `kthread_run_worker()` is an alias to `kthread_create_worker()`. The DRM drivers generic code started to replace `kthread_create_worker()` by `kthread_run_worker()` as is in Linux 6.14. Reviewed by: bz Sponsored by: The FreeBSD Foundation Differential Revision: https://reviews.freebsd.org/D57701
linuxkpi: Update `struct vm_area_struct` when an existing mapping is extended
With Mesa 26 and DRM drivers in Linux 6.13, userspace will try to extend
an existing mmap, at last push the end address further.
Before this change, the mapping was not updated, but userspace would try
to access a page after the initial end address, leading to a panic
triggered by the following assertion in `vm_fault_populate()`:
MPASS(fs->first_pindex <= pager_last);
Reviewed by: bz
Sponsored by: The FreeBSD Foundation
Differential Revision: https://reviews.freebsd.org/D58195
linuxkpi: Add `split_page()` This function is supposed to split large pages into an array of `PAGE_SIZE`-sized pages, with correct refcounting. This is apparently used to allow some drivers to free a part of a large page only. I don't think we use large pages in linuxkpi. Therefore, this new function is curently a no-op. The DRM drivers TTM memory manager started to use it in Linux 6.15. Reviewed by: bz Sponsored by: The FreeBSD Foundation Differential Revision: https://reviews.freebsd.org/D58762
linux: implement pkey_alloc, pkey_free and pkey_mprotect Bridge the Linux memory protection key syscalls to FreeBSD's native MPK support instead of returning ENOSYS. Modern Linux software probes these at startup: Chromium-based browsers (found via www/linux-brave) use protection keys for V8's heap and JIT sandboxing, and glibc >= 2.27 exposes the full API. pkey_alloc() allocates from a per-process bitmap kept in the process emuldata (key 0 implicitly allocated, matching Linux's mm_pkey_allocation_map; ENOSPC once keys 1..15 are exhausted or when PKU is absent, as Linux returns on such hardware) and applies the requested initial access rights to the calling thread's PKRU, located in the XSAVE area via xsave_area_offset(). pkey_free() is bookkeeping only: as on Linux, freeing neither untags pages nor updates PKRU. pkey_mprotect() performs the protection change and tags the range through amd64_pkru_update(), factored out of sysarch(2)'s AMD64_SET_PKRU/AMD64_CLEAR_PKRU implementation so that both share the same argument checking and map read lock synchronization with a parallel pmap_vmspace_copy() on fork; tags die with the mapping, matching Linux VMA semantics. A pkey of -1 degrades to plain mprotect. The allocation map is inherited on fork and reset on exec. At exec the Linux sysvecs initialize PKRU to 0x55555554, Linux's init_pkru default (access disabled for keys 1..15), so memory tagged with a not yet allocated key is inaccessible to threads that were never granted rights -- the property V8's thread isolation relies on. Setting PKRU at exec initializes the user FPU state slightly earlier than the lazy first-use path; the state would be initialized moments later in rtld/libc startup regardless. Protection key faults already deliver SEGV_PKUERR through the existing siginfo translation. The common code carries no architecture ifdefs. Machine-dependent state lives in struct linux_pemuldata_md, embedded in the process emuldata in the manner of struct mdthread, and common code calls per-arch lifecycle hooks (linux_pemuldata_init_md/_exec_md) and pkey back ends after performing the parameter validation Linux applies regardless of hardware support. On amd64 the implementation lives in sys/amd64/linux/linux_pkru.c, compiled into linux_common and serving both the 64-bit and 32-bit Linux ABIs. Elsewhere (arm64, i386) linux_emul_md.c provides stubs returning what Linux returns on hardware without protection keys (ENOSPC from pkey_alloc; pkey_mprotect with a pkey of -1 acts as plain mprotect), so applications take their normal no-PKU fallback instead of the ENOSYS path. PR: https://bugs.freebsd.org/bugzilla/show_bug.cgi?id=297427 MFC after: 1 month Reviewed by: kib Differential Revision: https://reviews.freebsd.org/D58782
linux: unbreak arm64 linux_emul_md.c after pkey syscalls linux_emul.h uses struct image_args without a file-scope forward declaration. The new arm64 (and i386) stubs include that header without imgact.h first, which fails the build under -Werror. Include it the same way linux_pkru.c already does, and declare the type next to struct image_params so the header is self-contained. Reported by: tuexen Fixes: https://cgit.freebsd.org/src/commit/?id=bdb561843e86 MFC after: 1 month
The K1 user manual lists three separate SDHCI devices but the upstream DTS file only defines the eMMC device. Create a temporary overlay to allow the BananaPi-F3 to boot from the SD card until this is addressed upstream. Differential Revision: https://reviews.freebsd.org/D57177 Reviewed by: mhorne
sys: Add sys/ckdint.h We have a C23 stdckdint.h header for userspace, which provides checked addition, subtraction and multiplication. We lack similar helpers in the kernel, where they are regularly needed. Let's just adopt the C23 macros. For bonus points, I added a wrapper to ensure that ignored an return value is raised as an error by the compiler. Reviewed by: kib, emaste MFC after: 2 weeks Sponsored by: The FreeBSD Foundation Differential Revision: https://reviews.freebsd.org/D58773
tools/build: stage stdckdint.h's dependencies for non-FreeBSD hosts
37bd69d43c7 gave stdckdint.h two new includes, <sys/_visible.h> and
<sys/ckdint.h>. Neither reaches a non-FreeBSD host: _visible.h is
staged only under ${.MAKE.OS} == "FreeBSD" and ckdint.h is not staged at
all, so the libc bootstrap fails on reallocarray.o when cross-building
from macOS. Both headers are self-contained; stage them alongside
stdckdint.h.
Fixes: https://cgit.freebsd.org/src/commit/?id=37bd69d43c7 ("sys: Add sys/ckdint.h")
Reviewed by: rpaulo, markj
Sponsored by: The FreeBSD Foundation
Differential Revision: https://reviews.freebsd.org/D58943
This serves to demonstrate some usage of the ckdint.h helpers. The new version also generates better machine code on amd64 and arm64. Reviewed by: kib, emaste MFC after: 2 weeks Sponsored by: The FreeBSD Foundation
Since the kernel environment has its own dependencies, lurking at the end of the SI_SUB_KMEM sequence appeared to be fragile. Provide own subsystem for it. The init_dynamic_kenv() goes SI_ORDER_FIRST, and two modules that depend on it go SI_ORDER_ANY. PR: https://bugs.freebsd.org/bugzilla/show_bug.cgi?id=297492 Reviewed by: imp, markj, emaste Differential Revision: https://reviews.freebsd.org/D58836
drm-kmod already implements the dma-buf and sync_file ioctls, but linux_ioctl.c had no handler group for the 'b' and '>' magic bytes, so the requests never reached it and returned EINVAL from linux_ioctl_fallback(). Route the commands drm-kmod services to sys_ioctl(), translating the direction bits with SETDIR(); everything else still falls through to the fallback and keeps getting named in dmesg. Approved-by: adrian Accepted-by: dumbbell Signed-off-by: Nick Price <nprice@FreeBSD.org> (cherry picked from commit d6a7e89504af337413af39fd121026f512c0a35d)
shmfd: consistently return size in 512 byte blocks for fstat(2) st_blocks This is ABI-breaking change that could be considered as the bug fix. Requested by: David Timber <dxdt@dev.snart.me> Reviewed by: emaste, markj Sponsored by: The FreeBSD Foundation Relnotes: yes MFC after: 1 week Differential revision: https://reviews.freebsd.org/D58942
tests/sys/posixshm/posixshm_test.c::accounting fix after st_size changes st_blksize is defined by POSIX as the 'preferred I/O block size for this object.' It is wrong to use st_blksize as the unit for st_blocks and expect it to be equal to the object size regardless of the change of st_blksize. Fixes: https://cgit.freebsd.org/src/commit/?id=3a1bf59d195c ("shmfd: consistently return size in 512 byte blocks for fstat(2) st_blocks") Reviewed by: markj Sponsored by: The FreeBSD Foundation MFC after: 1 week Differential revision: https://reviews.freebsd.org/D59014
Add more defines, structures, sort struct field types, add inline functions (partially implemented) all needed for the upcoming espwl(4) wireless driver. MFC after: 3 days
Fill in more (lvif) wdev details and add it to the list under the wiphy struct so that iterators at least work and find the (one) device. This is needed for the upcoming espwl(4) driver. misc: add WPI-SMS4 to the list of cipher suits (we won't support it but at least print the name). MFC after: 3 days
Convert IFT_BRIDGE and IFT_L2VLAN to ARPHRD_ETHER, and IFT_LOOP to ARPHRD_LOOPBACK in linux netlink. Also, add ARPHRD_SIT and convert IFT_STF to it. Reviewed by: kfv Differential Revision: https://reviews.freebsd.org/D58573
Noticed by: mckusick Fixes: https://cgit.freebsd.org/src/commit/?id=6563dcb6b1f57e51db63854f3774b52e672232ed
Reviewed by: bz Sponsored by: NVidia networking MFC after: 1 week Differential revision: https://reviews.freebsd.org/D58941
The scheduler selection interface uses names that are reserved words in C++, causing problems for downstream projects that use C++ in the kernel. Work around this by hiding the interface from C++ compilers until we can come up with a better solution. Fixes: https://cgit.freebsd.org/src/commit/?id=ce38acee8d0b ("Add kern/sched_shim.c") MFC after: 1 week Sponsored by: Klara, Inc. Sponsored by: NetApp, Inc. Reviewed by: siderop1_netapp.com, imp, kib Differential Revision: https://reviews.freebsd.org/D58991
6e474d8e3981 ("staging: vchiq_shim: avoid code duplication") refactors
some code which makes applying the subsequent patch easier.
49bec49fd7f2 ("staging: vc04_services: remove vchiq_copy_from_user")
addresses a user-triggerable integer overflow via the
VCHIQ_IOC_QUEUE_MESSAGE ioctl on /dev/vchiq (which has mode 0600 by
default). It also addresses insufficient validation of user-controlled
addresses in vchiq_copy_from_user().
Update the bcm2835_audio driver to follow the change to
vchi_msg_queue().
Reported by: Vicki Pfau
Reviewed by: Abdelkader Boudih <freebsd@seuros.com>
Tested by: Abdelkader Boudih <freebsd@seuros.com>
Tested by: Marco Devesas Campos <devesas.campos@gmail.com>
MFC after: 2 weeks
Sponsored by: The FreeBSD Foundation
Differential Revision: https://reviews.freebsd.org/D58889
The final consumer of this was OpenZFS, fixed in ffaea0831973 (thanks mav@). That change has been present in all active OpenZFS release branches for at least 6 months. These can finally be retired. Reviewed by: mav MFC after: 3 days Sponsored by: The FreeBSD Foundation Differential Revision: https://reviews.freebsd.org/D55201
The need for recursion applies (in somewhat different form) to both nullfs and unionfs, and would likely apply to any other hypothetical stacked filesystem as well. Reviewed by: kib, markj Differential Revision: https://reviews.freebsd.org/D58858
When a process execve()s, pmc_process_exec() is supposed to evaluate whether the new image is setuid/setgid and if so, whether to detach PMCs. This was handled by pmc_can_attach(), which is effectively an open-coded copy of cr_xids_subset(). Unfortunately, the test of the result of this function was inverted, with the result that we'd detach PMCs only if the predicate said it was okay to do so. It appears the bug has always been there; it seems the intent was to return 0 on "success", i.e., it is okay to attach the PMCs, much like p_candebug(). Commits 1c3c698ba4c4 and 1c40b15971f0 obscured this a bit. I think this check is trying to be too clever. Let's make it simpler: simply do not attach PMCs unless the owner is privileged. This is how, e.g., ktrace works. I do not think it's worth trying to be more sophisticated than this unless we can generalize the policy in a way that's applicable to other subsystems. Also fix a bug at the end of pmc_process_exec(): pmc_detach_one_process() will call pmc_remove_process_descriptor() for us. Approved by: so Security: FreeBSD-SA-26:56.hwpmc Security: CVE-2026-58089 Reported by: netchild Sponsored by: The FreeBSD Foundation Differential Revision: https://reviews.freebsd.org/D59102
This helper wasn't updated in commit be1f7435ef21, so in reality it was testing whether "gid" is the first supplemental group. If a user doesn't belong to a supplementary group, then it's testing an uninitialized slot; since ucreds are allocated with M_ZERO, this typically means that we're testing gid == 0. group_is_primary() has exactly one use, in mac_do. There, it's used to determine whether the requested primary GID can be used in a setcred(2) call when the ruleset does not explicitly specify a target primary GID. I believe this is mostly exploitable by daemons which have explicitly dropped privileges and called setgroups(0, NULL); logged in users will have a non-empty supplementary group list by virtue of having gone through initgroups(3). Fix group_is_primary(), and add a regression test. Approved by: so Security: FreeBSD-SA-26:59.mac_do Security: CVE-2026-58092 Reported by: Hazley Samsudin of GovTech CSG Fixes: https://cgit.freebsd.org/src/commit/?id=be1f7435ef21 ("kern: start tracking cr_gid outside of cr_groups[]") Reviewed by: olce, kevans Sponsored by: The FreeBSD Foundation Differential Revision: https://reviews.freebsd.org/D59051
The TIOCSCTTY ioctl handler drops the tty lock in order to acquire the proctree relock. After relocking the tty, it did not revalidate the tty state, and it could end up linking a doomed tty to the calling process' session. This race can be exploited to escalate privileges. TIOCSPGRP has a similar race, fix that too. Approved by: so Security: FreeBSD-SA-26:62.tty Security: CVE-2026-58093 Reported by: tsune of GMO Cybersecurity by Ierae, Inc. working with TrendAI Zero Day Initiative Reviewed by: kib Sponsored by: The FreeBSD Foundation Differential Revision: https://reviews.freebsd.org/D59126
The check for whether shm_lp_psind was assigned was unlocked. This race can be exploited to create an object with psind==2 but with only pagesizes[1] worth of pages populated. This in turn can be used to escalate privileges. Fix this by acquiring the rangelock earlier. In shm_mmap_large(), assert that we hold the rangelock. In shm_write(), annotate an unlocked load of shm_lp_psind. Approved by: so Security: FreeBSD-SA-26:63.posixshm Security: CVE-2026-58094 Reported by: tsune of GMO Cybersecurity by Ierae, Inc. working with TrendAI Zero Day Initiative Reviewed by: kib Sponsored by: The FreeBSD Foundation Differential Revision: https://reviews.freebsd.org/D59104
Define it everywhere that wants COMPAT_FREEBSD14. Reviewed by: imp, kib, emaste Sponsored by: OPNsense Sponsored by: Klara, Inc. Differential Revision: https://reviews.freebsd.org/D59017
Python scripts which use divert sockets no longer work after commit e967a2a03677; even if one patches socket() calls, getaddrlen() doesn't work on divert sockets, needed to use recvfrom(). Restore compatibility when COMPAT_FREEBSD15 is defined. Reviewed by: kib Sponsored by: OPNsense Sponsored by: Klara, Inc. Differential Revision: https://reviews.freebsd.org/D59018
If all member devices report the same rotation rate, pass it up.
Commit 08063e9f98 renamed cs_cpu to csr_cpu.
Finish reviewing all callers for lkpi_80211_mo_link_info_changed() and lkpi_80211_mo_bss_info_changed(), which are called from lkpi_bss_info_change() only. Add the lockdep_assert_wiphy() to lkpi_bss_info_change() and make sure all callers are holding the wiphy lock. PR: https://bugs.freebsd.org/bugzilla/show_bug.cgi?id=297228 Sponsored by: The FreeBSD Foundation MFC after: 3 days
Move the functions vtnet_rxq_csum() and vtnet_txq_offload() and the subfunctions they call from if_vtnet.c to virtio_net.h. This allows us to call these functions from if_tuntap.c and if_ptnet.c. virtio_net.h already contained a copy of these functions, but a copy of an outdated version. The functions evolved in if_vtnet.c. In if_vtnet.c, the copy has never been used because it increments counters in their own functions. This patch removes the outdated copy from virtio_net.h and moves the new version of the functions from if_vtnet.c to virtio_net.h. if_tuntap.c, if_ptnet.c, and if_vtnet.c just call these functions, and if_vtnet.c increments its counters depending on the return value. Reviewed by: tuexen MFC after: 1 month MFC to: stable/15 Differential Revision: https://reviews.freebsd.org/D57299
Adds missing structs symbols for V4L2. video(4) capture devices do not crop, expose menu controls or support overlay, and return ENOTTY for the new ioctls. Applications enumerate these unconditionally and degrade gracefully at run time, but fail to build when the declarations are missing. Reviewed by: adrian Differential Revision: https://reviews.freebsd.org/D59203
Reviewed by: adrian Differential Revision: https://reviews.freebsd.org/D59208
After rangelocks were reimplemented, _rangelock_cookie_assert() became a stub. Re-provide an implementation. Reviewed by: kib MFC after: 1 week Sponsored by: The FreeBSD Foundation Differential Revision: https://reviews.freebsd.org/D59222
When connecting a unix domain stream socket, we 1. look up the peer (listening) socket, 2. allocate a new socket 3. add the new socket to the listening socket's queue Prior to commit 26147c51546e, this sequence of operations was synchronized by a pool mutex, also acquired in uipc_close(). After commit 26147c51546e, we drop the vnode pool lock immediately after finding the peer socket via a filesystem lookup. This creates a window where it's possible for a connection to add a new socket to the listening queue after the listening queue has been aborted. Fix the race by restoring the old behaviour of holding the pool lock across the solisten_enqueue() call. This is a bit ugly since we need to pass a mutex lock and a vnode through a couple of layers, but it seems like a low-risk solution. Alternately we could add some flag to the listening socket which indicates that no new connections are to be accepted, but I think this will require some changes to the generic socket code. Reported by: pho Fixes: https://cgit.freebsd.org/src/commit/?id=26147c51546e ("unix: pin the pathname peer by reference across the connect") Reviewed by: olce, kib, John Ericson <John.Ericson@Obsidian.Systems> Differential Revision: https://reviews.freebsd.org/D59201
Reviewed by: markj Sponsored by: The FreeBSD Foundation MFC after: 1 week Differential revision: https://reviews.freebsd.org/D58094
The capability will allow the ptrace(2) on the procdesc. Reviewed by: markj Sponsored by: The FreeBSD Foundation MFC after: 1 week Differential revision: https://reviews.freebsd.org/D58586
If the flag is not specified, the process descriptor returned by either pdfork(2) or pdopenpid(2) has the CAP_PTRACE capability disabled. Reviewed by: markj Sponsored by: The FreeBSD Foundation MFC after: 1 week Differential revision: https://reviews.freebsd.org/D58586
Reviewed by: markj Sponsored by: The FreeBSD Foundation MFC after: 1 week Differential revision: https://reviews.freebsd.org/D58586
The code to handle copyin and copyout of the structured parameters is moved into the helpers. Reviewed by: markj Sponsored by: The FreeBSD Foundation MFC after: 1 week Differential revision: https://reviews.freebsd.org/D58586
Reviewed by: markj Sponsored by: The FreeBSD Foundation MFC after: 1 week Differential revision: https://reviews.freebsd.org/D58586
Reviewed by: markj Sponsored by: The FreeBSD Foundation MFC after: 1 week Differential revision: https://reviews.freebsd.org/D59113
The function defines the policy for allowing to open a pid. Reviewed by: markj Sponsored by: The FreeBSD Foundation MFC after: 1 week Differential revision: https://reviews.freebsd.org/D58989
The pdopenpid(2) syscall is allowed in capability mode. Add the chicken switch security.bsd.ptrace_in_cap_mode, which disables it without reboot, if needed. The descriptor passed to pdptrace(2) must have the CAP_PTRACE capability enabled. This capability is not enabled by default by pdfork()/pdopenpid(), and the calls do not return a procdesc suitable for debugging. The opening code must prepare for debugging in advance by passing the PD_PTRACE_CAP flag to pdfork()/pdopenpid(). For ptrace(2), allow PT_CLEARSTEP and PT_GET_CHILDREN for the current thread and process in cap mode as well. Reviewed by: markj Sponsored by: The FreeBSD Foundation MFC after: 1 week Differential revision: https://reviews.freebsd.org/D58989
The pdopenpid() syscall is allowed to open processes which are either direct children of the caller, or are debuggees already attached to the calling process. This is reasonable because we could have controlled the child on fork anyway. The procdesc-less debuggee can legitimately appear due to ptrace FOLLOW-FORK mode. Reviewed by: markj Sponsored by: The FreeBSD Foundation MFC after: 1 week Differential revision: https://reviews.freebsd.org/D58989
Originally gstripe created a separate child I/O for every accessed stripe, making it very inefficient for small stripe sizes. Later introduced "fast" mode reduced that count for read/write requests by copying the data to/from temporary contiguous buffers, wasting memory bandwidth and CPU time. This commit implements alternative method, utilizing unmapped I/O mechanism to assemble children I/Os from pages of the original I/O, avoiding any copying. This method though has some limitations, such as stripe size can not be smaller than CPU page size, or buffer and offset page phases should match (may be page aligned, but not necessarily). But those limitations are not an issue in many cases, since ZFS, for example, can often align its buffers (BTW, dd doesn't). Plus, unlike "fast" method, this one can receive (and even prefers) unmapped I/Os. While there, re-implement also BIO_DELETE. Since they don't have any data, there is no any reason to create more than one child I/O per disk. It also dramatically improves performance there. And for dessert, add rotation rate attribute support, reporting one if all the children report the same. ZFS is using it more and more.
Reviewed by: markj Tested by: pho Sponsored by: The FreeBSD Foundation MFC after: 1 week Differential revision: https://reviews.freebsd.org/D59132
tty: add tty_wait_proctree(9) Reviewed by: markj Tested by: pho Sponsored by: The FreeBSD Foundation MFC after: 1 week Differential revision: https://reviews.freebsd.org/D59132
tty: make tty_wait_background() aware of proctree_lock ownership Reviewed by: markj Tested by: pho Sponsored by: The FreeBSD Foundation MFC after: 1 week Differential revision: https://reviews.freebsd.org/D59132
tty: gracefully handle proctree_lock locking Instead of relocking tty to get the proctree_lock and experiencing the race due to the relock, take the proctree_lock in advance for ioctl commands that need it. The affected commands, TIOCNOTTY, TIOCSCTTY, and TIOCSPGRP, must not be overridden by the specific tty drivers, so the common handling is cleaner. Reviewed by: kevans (previous version), markj Tested by: pho Sponsored by: The FreeBSD Foundation MFC after: 1 week Differential revision: https://reviews.freebsd.org/D59132
If all member devices report the same rotation rate, pass it up.
If all member devices report the same rotation rate, pass it up.
This patch adds assorted bits needed by the nfsclrdma.ko module that implements client side NFS over RDMA. It should not affect non-RDMA operation. Some additional glue is needed for the nfsclrdma.ko module within the NFS code. That will be added as a separate commit. I've specified a long MFC, since the module still requires extensive testing and, hopefully, a review. MFC after: 3 months
There seems to be another possible race with net80211 state machine changing the bss from under us (another lvif_bss_synched case). Just do the != NULL check to avoid a NULL pointer deref in ieee80211_ratectl_rate(). (bz extended the original comment and wrote the commit message). Sponosred by: The FreeBSD Foundation (commit) PR: https://bugs.freebsd.org/bugzilla/show_bug.cgi?id=297184 MFC after: 3 days
This compatibility alias for kern.elf{32,64}.fallback_brand has been
marked for removal since FreeBSD 6.0.
Signed-off-by: Christos Longros <chris.longros@gmail.com>
Reviewed by: kib, emaste
Differential Revision: https://reviews.freebsd.org/D56085
lkpi_cfg80211_calculate_bitrate_vht() was constantly showing up in my debug traces as a TODO with rtw89 so I went ahead and implemented the HT and VHT versions. Realtek seems to limit amsdu sizes based on the value and ask for it whether needed or not. Sponsored by: The FreeBSD Foundation MFC after: 3 days
umtxq_hash() multiplies the key by 0x9E370001 and keeps the high bits. That constant is 0x9E37 * 2^16 + 1, so it degenerates for keys whose spacing carries trailing zero bits: at a 64 KiB stride it puts 128 of 512 parked waiters onto a single chain mutex, and at 16 KiB and up it uses only a handful of the 512 chains. Base-system consumers never hit this because libthr places its own wait words 128 bytes apart, but a Linux-ABI runtime waiting on addresses it allocates itself lands squarely on the floor. Switch to 0x61C88647, which leaves at most 3 waiters per chain at the same stride; Linux made this exact change in 2016, after judging the sparse constants "actively bad for hashing". Approved by: adrian (mentor) Reviewed by: kib, adrian, emaste Differential Revision: https://reviews.freebsd.org/D58337 Signed-off-by: Nick Price <nprice@FreeBSD.org>
to not leak information about unused pids or system processes' pids. Reviewed by: markj Fixes: https://cgit.freebsd.org/src/commit/?id=73c92a978cce ("pdopenpid(2): allow in capability mode with restrictions") Sponsored by: The FreeBSD Foundation MFC after: 1 week Differential revision: https://reviews.freebsd.org/D59252
There were a couple of problems detected w.r.t. the "glue" for the nfsclrdma.ko module. - When the NFS server has a small reply for a read, it can choose to not use the reduction chunk (separate memory area for the read data). I did not realize this was the case. - There was a bug in rpc_copy_uio_pages() function that caused intermittent crashes in memcpy(). This patch fixes the above cases. It uses M_PROTO6 to mark that an RPC reply has used a reduction chunk, so that read can handle it correctly. Read also now provides a reduction chunk for all read sizes, since the worst case for the rest of the read RPC reply is close to the 1024 byte limit. (NFSv4 uses strings instead of uid/gid in the attributes and these name strings can be rather large.) I wanted to get the "glue" into main so that others could test the module more easily. Avaliability of the module will be announced on freebsd-current@ soon. It should not affect non-RDMA operation. I've specified a long MFC, since the module still requires extensive testing and, hopefully, a review. MFC after: 3 months Fixes: https://cgit.freebsd.org/src/commit/?id=884ee8d6c9b4 ("nfscl: Add some glue for client side NFS over RDMA")
- Add unmapped I/O support. The only case when the code needs data access is BIO_READ returning zeroes for unallocated space. - Add BIO_FLUSH support. Just send it to all allocated components. - Add BIO_DELETE support. While current design does not allow freeing allocated blocks, at least pass it to underlying providers. - Add direct I/O completion support. - Add rotation rate reporting. - Fix few minor issues.
Pushed using the RTL8723BU. Reviewed by: ziaee, avos, adrian Relnotes: yes Differential Revision: https://reviews.freebsd.org/D59205
Reviewed by: markj Sponsored by: The FreeBSD Foundation MFC after: 1 week Differential revision: https://reviews.freebsd.org/D59317
The 16-byte "device_name" field was not zero-filled, so could contain uninitialized stack data. Zero the whole struct, as that's the prevailing pattern for this kind of conversion code, and it's more robust in the face of future revisions to struct devstat. Reviewed by: olce, kib Reported by: Reo Shiseki Fixes: https://cgit.freebsd.org/src/commit/?id=a11d132f6c62 ("devstat: Provide 32-bit compatibility") MFC after: 3 days Sponsored by: The FreeBSD Foundation Differential Revision: https://reviews.freebsd.org/D59309
Make the witness LOCK_CHILDCOUNT a configurable kernel option. On machines with a very high core count the default value is too low, leading to witness exhaustion after boot. Relnotes: yes Reviewed by: kib, ziaee Signed-off-by: Kajetan Puchalski <kajetan.puchalski@arm.com> Closes: https://github.com/freebsd/freebsd-src/pull/2398
Add iri_rcv_tstmp to if_rxd_info so an isc_rxd_pkt_get() driver can report a hardware RX timestamp. Copy it into m_pkthdr.rcv_tstmp, reusing the generic mbuf timestamp path. Widen iri_flags from uint8_t to uint32_t and define the flags drivers may supply. Mask the flags before copying them into the mbuf so no other mbuf state can leak through the driver callback. Place the timestamp next to iri_frags to avoid an alignment hole, and document its nanoseconds-since-boot representation and validity flags. Bump __FreeBSD_version because changing if_rxd_info breaks KBI. Reviewed by: gallatin Signed-off-by: Sreekanth Reddy <sreekanth.reddy@broadcom.com> Differential Revision: https://reviews.freebsd.org/D58638
A couple of additional fixes for the NFS client side RDMA glue: - For Readdirplus, the reply needs to be a large chunk, so set M_PROTO9 instead of M_PROTO8. - The nfsclrdma.ko module now uses xprt_rdma_unmap_chunk() instead of xprt_rdma_rekey_chunk(). Hopefully, this is it for the NFS over RDMA client glue changes. MFC after: 3 months Fixes: https://cgit.freebsd.org/src/commit/?id=884ee8d6c9b4 ("nfscl: Add some glue for client side NFS over RDMA")
Use an unprivileged load to access user memory from dtrace_copy, which is running in SVC mode. Abort the loop if the load is trapped, as it is useless, hence wasteful, to keep faulting on successive addresses. I believe that this de-pessimization should also be done on aarch64 and riscv. PR: https://bugs.freebsd.org/bugzilla/show_bug.cgi?id=298064 MFC after: 1 month Reviewed by: markj Differential Revision: https://reviews.freebsd.org/D59280
Fix the constant case label to properly handle translation faults caused by DTrace probes. Alignment errors are not expected to be generated, so stop handling them. While at it, correct an amd64-specific comment and add a comment regarding the missing faulting address which could be addressed by a later improvement. PR: https://bugs.freebsd.org/bugzilla/show_bug.cgi?id=298064 MFC after: 1 month Reviewed by: markj Differential Revision: https://reviews.freebsd.org/D59281
On ll/sc architectures casueword32() may report a spurious
store-conditional failure (reservation lost to an interrupt, preemption,
or another CPU touching the same reservation granule), and this is
indistinguishable from a genuine comparison mismatch: both return 1.
That is intentional since D20772 and documented in casueword(9) ("The
store can fail on load-linked/store-conditional architectures."), so
callers must cope.
do_lock_normal() does not fully cope. When the initial
UMUTEX_UNOWNED -> id acquire CAS fails spuriously, the observed owner is
still UMUTEX_UNOWNED, so neither the UMUTEX_CONTESTED branch nor the
real-owner case applies, and execution falls through past the "rv == 1
but not contested, likely store failure" comment into the sleep path.
There, the contested-bit CAS (expecting the observed owner, i.e.
UMUTEX_UNOWNED) succeeds because the mutex really is unowned, stamping
m_owner = UMUTEX_CONTESTED with no owner tid, and the thread sleeps on
"umtxn" forever: nobody owns the mutex, so no unlock and no wakeup ever
arrive. In _UMUTEX_TRY mode the same situation returns a false EBUSY
for a free mutex.
Treat an observed owner of UMUTEX_UNOWNED like UMUTEX_CONTESTED: try to
acquire the mutex, setting the contested bit, instead of falling through
to the sleep path. rv == 1 with the observed value equal to the
expected value can only mean a spurious store failure, so the mutex is
free. If the acquire CAS fails again, the outer loop restarts and
re-evaluates ownership. The contested bit set with no waiters present
only costs the matching unlock one trip through the kernel.
This was hit in practice on powerpc64le (POWER9): the Swift runtime's
Synchronization.Mutex issues _umtx_op(UMTX_OP_MUTEX_LOCK) directly with
no userspace fast path, so an uncontended lock of an unowned mutex runs
the kernel CAS exactly where a spurious failure deadlocks
(single-threaded process parked on "umtxn" with m_owner == 0x80000000,
observed as Foundation.Process.run() hanging). libthr mostly masks the
bug because pthread_mutex_lock() enters the kernel only when there is a
real owner that will eventually issue a wakeup.
The mechanism was confirmed with an experimental powerpc kernel that
instead retried the ll/sc sequence inside casueword32()/casueword();
that also eliminated the hang, but is not proposed here since the
single-attempt semantics of casueword(9) are intentional.
Reviewed by: kib
MFC after: 2 weeks
Differential Revision: https://reviews.freebsd.org/D59338
bufinit() inserts newly initialized bufs into the QUEUE_EMPTY queue, at which point they haven't yet been assigned a domain. Thus, bufdomain() returns &bdomain[-1], which trips the array-bounds sanitizer. This is harmless since we don't use the result in that case, but let's avoid the invalid access to begin with. This is sufficient to let an amd64 kernel boot to a login prompt with -fsanitize=array-bounds configured. Reported by: Andrew Griffiths <andrew@calif.io> Reviewed by: rlibby, kib MFC after: 1 week Differential Revision: https://reviews.freebsd.org/D59381
But only if debug.bootverbose=1. MFC after: 2 weeks Sponsored by: ConnectWise
In dma_sync_single_for_cpu(), the DMA_BIDIRECTIONAL direction currently performs BUS_DMASYNC_POSTREAD followed by BUS_DMASYNC_PREREAD. This patch corrects the mapping to use BUS_DMASYNC_POSTREAD | BUS_DMASYNC_POSTWRITE. When ownership of the DMA area is transferred to the CPU, we must assume the previous device access was bidirectional. Both POST operations are necessary to ensure the CPU sees a consistent view of memory after potential device reads and writes. A PREREAD is unnecessary here because the device will no longer access the memory since ownership has been transferred to the CPU. Conversely, for dma_sync_single_for_device(), ownership is being transferred back to the hardware. The buffer must be prepared for potential bidirectional access by the device, requiring BUS_DMASYNC_PREREAD | BUS_DMASYNC_PREWRITE. PR: https://bugs.freebsd.org/bugzilla/show_bug.cgi?id=293381, https://bugs.freebsd.org/bugzilla/show_bug.cgi?id=297155 Reported by: Zishun Yi <zishun.yi.dev@gmail.com> Reported by: Kim Shrier (fbsdbugs westryn.net) Fixes: https://cgit.freebsd.org/src/commit/?id=95edb10b47fc ("LinuxKPI: implement dma_sync_single_for_*, apply to (un)map single/sg") Signed-off-by: Zishun Yi <zishun.yi.dev@gmail.com> Reviewed by: aokblast, bz Differential Revision: https://reviews.freebsd.org/D55497
Both kern_renameat() and kern_dup() had arguments named `new`. Rename their arguments to match their respecitve manual pages. Sponsored by: Klara, Inc. Sponsored by: NetApp, Inc. Reviewed by: kib Differential Revision: https://reviews.freebsd.org/D59386
The entropy will be collected in intr_event_schedule_thread() for event handlers that are marked as entropy sources. The quality of entropy provided by SWI wasn't great anyway. Reviewed by: cem, ngie Differential Revision: https://reviews.freebsd.org/D58829
Renumbered the camera class to match the reference ABI and added the constants ported applications expect: Reviewed by: thierry Differential Revision: https://reviews.freebsd.org/D59459
The camera control IDs in videoio.h changed numeric values; ports consuming V4L2 camera controls need to be rebuilt. Reviewed by: thierry Differential Revision: https://reviews.freebsd.org/D59460
Implement dma_sync_sg_for_{cpu, device}() and
dma_sync_sgtable_for_device().
These functions are useful for my GSoC 2026 project, udmabuf.
Reviewed by: bz
MFC after: 3 days
Differential Revision: https://reviews.freebsd.org/D57766
According to Linux documentation the nents argument to dma_unmap_sg() must be the number one passed in, not the number of DMA addresses. In LinuxKPI this means orig_nents and not nents, so adjust this. Given nents and orig_nents should always be the same in LinuxKPI, this should only be a NOP for correctness. Reviewed by: bz, aokblast (LGTM) MFC after: 3 days Differential Revision: https://reviews.freebsd.org/D57842
Bump for the new ifnet and iflib VF-status provider KBI and the rtnetlink VF-status interface.
nfsd: Clean up the "glue" for the nfsrdma.ko module Move svc_reg() calls into a helper function so that the nfsrdma.ko can call that. Create a new nfs_extern.h as a place to put the nfs stuff that server side NFS over RDMA needs to access, with a prototype for the helper function and a couple of definitions that probably shouldn't be in svc.h. No semantics change. MFC after: 3 months Fixes: https://cgit.freebsd.org/src/commit/?id=7144a1d58c5c ("nfsd: Add glue for the nfsrdma.ko module")
nfsd: Add nfs_extern.h Oops, forgot to add the new .h file. MFC after: 3 months Fixes: https://cgit.freebsd.org/src/commit/?id=407d7177057d ("nfsd: Clean up the "glue" for the nfsrdma.ko module")
The vhci portion of the driver relies on usb(4) for both constants and functionality. Add an explicit dependency on usb(4) in `files.amd64` so it builds properly 100% of the time in a static kernel config. Reviewed by: seuros Differential Revision: https://reviews.freebsd.org/D59469
Since the vp_crossmp vnode can leak into vn_vptocnp() calls due to nullfs file mounting, not all lock requests are non-sleeping. The requirement for the crossmp locking is that all lock requests should be shared. Then, it does not matter if the requests allow sleeping, since all locks are shared. Also, check the lock type by correctly masking it with LK_TYPE_MASK. Reported and tested by: pho Reviewed by: jah, markj Sponsored by: The FreeBSD Foundation MFC after: 1 week Differential revision: https://reviews.freebsd.org/D59468
Reviewed by: markj Sponsored by: The FreeBSD Foundation MFC after: 1 week Differential revision: https://reviews.freebsd.org/D59349
Before this change, the assertion would not catch a page index matching the end address, even though the end address of a mapping is outside of that mapping. While here, add another assertion to catch page indexes outside of the expected range earlier. Also, pagefaulting an address outside of the mapping range should return `VM_FAULT_SIGBUS` instead of panicing. Reviewed by: bz Sponsored by: The FreeBSD Foundation Differential Revision: https://reviews.freebsd.org/D58194
The function converts a number of seconds since Epoch to a `struct tm`. The implementation was copied from `contrib/tzcode/localtime.c`. The amdgpu DRM driver started to use it in Linux 6.15. Reviewed by: bz Sponsored by: The FreeBSD Foundation Differential Revision: https://reviews.freebsd.org/D58758
`hrtimer_setup()` replaces `hrtimer_init()` in Linux 6.15. The API and the role are the same, except: * it takes the `function` callback as argument * if `function` is NULL, it assigns a default callback At the same time, we hide `hrtimer_init()` if LINUXKPI_VERSION is greater than or equal to 61500 because it was dropped in Linux 6.15. Reviewed by: bz Sponsored by: The FreeBSD Foundation Differential Revision: https://reviews.freebsd.org/D58757
The amdkfd driver requires the `tgid` to be a part of the `task_struct`. This patch introduces the `tgid` member to `task_struct`. Reviewed by: bz Sponsored By: The FreeBSD Foundation Differential Revision: https://reviews.freebsd.org/D58228
This macro was moved from <drm/drm_util.h> to <linux/util_macros.h> in Linux 6.15. Submitted by: Sourojeet Adhikari Approved by: adrian, bz, emaste, seuros Sponsored by: The FreeBSD Foundation Differential Revision: https://reviews.freebsd.org/D57465
jiffies64_to_msecs is used by kfd_process.c in the amd/amdkfd driver from Linux kernel 6.12. Submitted by: Sourojeet Adhikari Sponsored by: The FreeBSD Foundation Differential Revision: https://reviews.freebsd.org/D58075
linuxkpi: Add device under parent, not under class
In `device_add()`, the function used to add the given device under its
class. This is used to build a sysctl tree. We ended up with devices or
"pseudo" devices (like the output connectors of a GPU). For example with
an output connector:
sysctl sys.class.drm.card0-DP-1
This device should be added under its parent if it has one. With this
fix, the same output connector is now:
sysctl sys.device.drmn1.card0.card0-DP-1
Reviewed by: bz
Sponsored by: The FreeBSD Foundation
Differential Revision: https://reviews.freebsd.org/D55175
linuxkpi: Add `pci_map_rom()` and `pci_unmap_rom()` They were already defined as macros in various places in DRM drivers, aliasing the `vga_pci_map_bios()` and `vga_pci_unmap_bios()` functions. Let's move them to linuxkpi and avoid copies everywhere. Because they use the `vga_pci` code internally, `pci_map_rom()` checks whether the given device is a video card. If it is not, it logs a "TODO" and returns NULL. Reviewed by: bz Sponsored by: The FreeBSD Foundation Differential Revision: https://reviews.freebsd.org/D57573
linuxkpi: Define `DEFINE_CLASS()` and `CLASS()` `DEFINE_CLASS()` is in fact named `LINUXKPI_DEFINE_CLASS()` because it conflicts with `DEFINE_CLASS()` defined in <sys/kobj.h>. This macro defines a type and a pair of constructor/destructor functions. They are to be used by `CLASS()`: this one declares a variable, initialise it with the constructor and set the `__cleanup()` attribute to call the destructor once the variable goes out of scope. The DRM drivers generic code started to use `CLASS()` in Linux 6.13. It requires the `fd` class to be defined in <linux/file.h>. Reviewed by: bz Sponsored by: The FreeBSD Foundation Differential Revision: https://reviews.freebsd.org/D57583
When DTrace catches a data abort exception it resumes execution on the next instruction. If a probe executes copyinto from unmapped memory, then dtrace_copy keeps faulting on successive bytes until the loop counter is exhausted. Detect the situation by witnessing the absence of zero-extension from an aborted unprivileged load. Reviewed by: markj MFC after: 2 weeks Differential Revision: https://reviews.freebsd.org/D59341
Userland jail managers hardcode jail parameter name strings; nothing in the headers has ever named them, even when sys/jail.h already has JAIL_META_PRIVATE/JAIL_META_SHARED for the metadata pair. Added JAIL_PARAM_* string constants for the static parameters registered by the base kernel. No functional change. MFC After: 1 week Reviewed by: jamie, adrian Differential Revision: https://reviews.freebsd.org/D59572
The implementation comes from `isqrt64()` in `sys/cam/cam_iosched.c`. It is modified to take an `unsigned long` and return an `unsigned int`. The i915 DRM driver started to use it in Linux 6.15. Reviewed by: bz Sponsored by: The FreeBSD Foundation Differential Revision: https://reviews.freebsd.org/D58251
When READ CAPACITY fails, we call cam_periph_invalidate() for some hopeless ASC/ASCQ combinations and then fall through to test the cache state. This results in excess commands to the abandoned device that may take a while to timeout and/or report confusing errors. Instead, call daprobedone() when the periph is invalid to terminate the probe. Fixes: https://cgit.freebsd.org/src/commit/?id=d78d04b17cb24 Sponsored by: Netflix MFC After: 1 week Differential Revision: https://reviews.freebsd.org/D58524
Apply de-morgan's law correctly to force only 4/2 to go to the invalidate stage for that ASC/ASCQ. Fixes: https://cgit.freebsd.org/src/commit/?id=d78d04b17cb24 Sponsored by: Netflix MFC After: 1 week Reviewed by: ali_mashtizadeh.com, mav Differential Revision: https://reviews.freebsd.org/D58525
To be able to create other kinds of devctl notifications that don't flow through cam_periph_devctl_notify, refactor out the header / trailers we do and the sbuf life cycle management. No functional change. Sponsored by: Netflix Differential Revision: https://reviews.freebsd.org/D58526
Report serial number, if available, otherwise report the path of the device invalidated. Sponsored by: Netflix Differential Revision: https://reviews.freebsd.org/D58527
cam::periph:error, cam::{,a,n,sd}da:error
Trace the error evolution for I/O errors the periph encounters
cam::periph:recovery, cam::{,a,n,sd}da:recovery
Trace the result of the error sorting
cam::periph:invalidate
Trace device invalidations
cam::periph:hold-boot
cam::periph:release-boot
cam::xpt:hold-boot
cam::xpt:release-boot
Trace boot holds and releases at different levels
cam::xpt:bus-register
Trace new busses added
These additions allow fairly complete instrumentation of the discovery
process at boot through dtrace's anonymous tracing. It would be mildly
better to have the boot holds have the periph, but dtrace sbts have to
be static at compile time.
Sponsored by: Netflix
Reviewed by: dteske
Differential Revision: https://reviews.freebsd.org/D58523
linux: implement prctl PR_{G,S}ET_THP_DISABLE
Since we don't have THP in FreeBSD, return that it's disabled. Linux
applications use this to get around the weird latency spikes THP in
Linux creates. Since FreeBSD superpages don't suffer from this issue,
just say it's disabled.
This removes the noisiest noise from the linux claude running on
FreeBSD.
Sponsored by: Netflix
linux: LINUX_PR_SET_THP_DISABLE all non-zero values are disable Linux treats any nonzero value as "disable", so accept them all. Fixes: https://cgit.freebsd.org/src/commit/?id=a8a6eac57091 Sponsored by: Netflix
Adding the `compat_ptr_ioctl` function for amdkfd support it's called by kfd_chardev.c Reviewed by: dumbbell Differential Revision: https://reviews.freebsd.org/D58076
Add support for hierarchical softc layout so that a leaf class and each of its base classes owns a private softc region inside a single allocation. device_get_softc_class(dev, cls) returns a pointer to the softc that belongs to the requested class. The classic device_get_softc() still returns the leaf softc and remains fully compatible with existing drivers. Existing drivers are unaffected; they simply obtain a slightly larger softc block when they inherit from base classes. MFC after: 2 months Reviewed by: kib Differential Revision: https://reviews.freebsd.org/D59115
Linux executes irq_work callbacks from an interrupt context that provides implicit RCU read-side protection. LinuxKPI dispatches these callbacks through a taskqueue, so provide equivalent protection explicitly. This prevents an RCU grace period from completing while an irq_work callback is still using an RCU-protected object. Suggested by: wulf Tested by: JustAnotherHumanBeing Signed-off-by: JustAnotherHumanBeing <oleglelchuk@gmail.com> Pull Request: https://github.com/freebsd/freebsd-src/pull/2385
bridge_pfil() pulled up min(m_pkthdr.len, max_protohdr) bytes. When the mapped head is shorter than that and followed by an unmapped (M_EXTPG) mbuf -- a sendfile(2) or KTLS segment from a member advertising IFCAP_MEXTPG -- m_pullup() ran into it and dereferenced a NULL mtod(), panicking the kernel. Pull up the Ethernet header first, and the SNAP/LLC header only for an 802.3 frame. This is similar to pf and ip_output(). m_pullup() and m_copyup() asserted only the first mbuf; assert inside both copy loops so the shape trips the check. Fixes: https://cgit.freebsd.org/src/commit/?id=c38abd64dbc1 ("if_epair: support IFCAP_MEXTPG") Suggested by: markj Reviewed by: markj, gallatin Assisted-by: Claude Code (Fable 5, Opus 5)
This should be it for a while, but there will be another cycle of "glue" updates. I just found out that I'll need to create an alternate code path that uses a contigmalloc() blob instead of scatter/gather of pages, since some NICs cannot do the scatter/gather of pages well. This commit should not affect non-RDMA behaviour. MFC after: 3 months Fixes: https://cgit.freebsd.org/src/commit/?id=884ee8d6c9b4 ("nfscl: Add some glue for client side NFS over RDMA")
This was a rather dumb miss on my part in commit 42442d7a6e.
LK_CANRECURSE is clearly needed in any case in which the covered vnode
is held exclusive across the call to VFS_ROOT(), regardless of whether
it was initially held exclusive or upgraded. The commit message for
that change also noted that unionfs lookup only worked without
LK_CANRECURSE due to a coincidence of the then-current unionfs
implementation. As it happens, said coincidence was recently removed
in commit b952606b4f ("unionfs_lock(): eliminate LK_CANRECURSE special-
case").
PR: https://bugs.freebsd.org/bugzilla/show_bug.cgi?id=298201
Reported by: olivier
Fixes: https://cgit.freebsd.org/src/commit/?id=42442d7a6e "Generalize the VV_CROSSLOCK logic in
vfs_lookup"
Reviewed by: kib, markj, pho
Tested by: pho
MFC after: 1 week
Differential Revision: https://reviews.freebsd.org/D59494
Align behavior between the Linux cancel_delayed_work_sync() function return value and the LinuxKPI equivalent. Linux cancel_delayed_work_sync() returns whether delayed work was pending, even if canceled before executing. This includes the case where the timer fired and work was queued but the callback had not yet started. The LinuxKPI version used the return value from taskqueue_cancel() as the return value of the public facing API, which inverted the behavior of two cases, violating the Linux API contract. Queued work which was removed before running would return false, and work whose callback was already executing would return true. Track the taskqueue pending count separately from the taskqueue_cancel() return value. Use the pending count for the public return value. Use taskqueue_cancel() return value only to decide whether state needs to be re-checked. This fixes behavior in consumers which use the return value of cancel_delayed_work_sync() to distinguish canceled pending work from work that was already running or idle. Signed-off-by: Ryan Fahy <ryan@rfahy.com> Reviewed by: wulf MFC after: 1 month Pull Request: https://github.com/freebsd/freebsd-src/pull/2268
Rename the sendfile() function to kern_sendfile(), expand the arguments previously passed in struct sendfile_args, and extend with two function pointer arguments to copy in the header/trailer structure and the create uio's for the header and trailer as required. Use this to allow the removal of freebsd32_do_sendfile() which was a nearly identical duplicate of sendfile() with attendant maintenance cost. Reviewed by: kib, markj Effort: CHERI upstreaming Sponsored by: Innovate UK Differential Revision: https://reviews.freebsd.org/D59034
The lock acquisition in g_raid3_ctl_insert() was improperly dropped a while ago, making it impossible to add or replace a device in an existing graid3. This went unnoticed because the tests are broken. MFC after: 1 week Fixes: https://cgit.freebsd.org/src/commit/?id=fcf69f3dbce6 ("Consistently use gctl_get_provider instead of home-grown variants.") Event: EuroBSDcon 2026 DevSummit Reviewed by: delphij Differential Revision: https://reviews.freebsd.org/D59563
On Linux `dma_length` field of `struct scatterlist` is present on the
arches where DMA mapping code is able to coalesce adjacent segments
of physical address space. It contains total length of coalesced
segments while `length` field contains non-coalesced length of each
segment. On other arches `dma_length` is aliased to `length` field with
`sg_dma_len` macro. As FreeBSD does not merge scatterlist segments it
do not have `dma_length` field. It is appered that at least i915kms
driver depends on existence of `dma_length` field.
Add the field and disable it by default. To enable add to Makefile
.if ${MACHINE_CPUARCH} == "i386" || ${MACHINE_CPUARCH} == "amd64" || \\
${MACHINE_CPUARCH} == "aarch64" || ${MACHINE_CPUARCH} == "powerpc"
CFLAGS+= -DCONFIG_NEED_SG_DMA_LENGTH
.endif
Reported by: Ryan Fahy
Reviewed by: bz
MFC after: 1 month
GHI: https://github.com/freebsd/drm-kmod/issues/315
Differential Revision: https://reviews.freebsd.org/D59630
The active LBA format's MS field was never examined. I/O to a metadata-formatted namespace carries neither interleaved metadata nor MPTR, so every command is malformed, yet the namespace attaches as a disk with the wrong sector size. Reviewed by: imp, adrian Differential Revision: https://reviews.freebsd.org/D59625
Yet again. I was trying to make the svc_vc_backchannel() operations do double duty and be used by the clnt_rdma.c code as well. It got too messy, so this reverts svc_vc.c back to its pre-glue form and adds the small changes needed to support a separate set of svc_rdma_backchannel_xxx() functions. This commit should not affect non-RDMA behaviour. MFC after: 3 months Fixes: https://cgit.freebsd.org/src/commit/?id=884ee8d6c9b4 ("nfscl: Add some glue for client side NFS over RDMA")
This fixes the build for kernels with INVARIANT_SUPPORT but without INVARIANTS. Fixes: https://cgit.freebsd.org/src/commit/?id=2e376cca379b ("rangelock: Reimplement _rangelock_cookie_assert()")
Add the missing locking to another two MO driver downcalls, as drivers always expect it (e.g. rtw89 by assertion). Add the lock and might_sleep assertions to the respective lkpi_80211_mo_* downcalls. PR: https://bugs.freebsd.org/bugzilla/show_bug.cgi?id=298417 Reviewed by: bz MFC after: 3 days Differential Revision: https://reviews.freebsd.org/D59608
The size of multiple embedded structs have changed and may lead to problems (pci_error_handlers in pci_driver, dev_pm_info in struct device). Allow these changes to be detected by bumping __FreeBSD_version. MFC after: 3 days
There are two parts that go together: * `X86_MATCH_VFM()` to declare a matching pattern * `x86_match_cpu()` to check if the current CPU matches one of the patterns in an array. The i915 DRM driver started to use this in Linux 6.14. Reviewed by: kib Sponsored by: The FreeBSD Foundation Differential Revision: https://reviews.freebsd.org/D57700
The generic if_qflush() only clears if_snd. Calling it directly for a transparent VF bypasses the driver's callback and its private software transmit queues. Add if_getqflushfn(), matching the existing callback accessors, and invoke the registered callback while hn_vf_lock protects the VF pointer. Flush an attached VF whenever transparent mode owns it, not only while the VF datapath is enabled. Disabling transmit during handoff does not discard packets already queued to the VF. This also covers the inactive interval after fallback. Non-transparent VFs remain independently owned and are not flushed by hn. Synthetic queue flushing is unchanged. MFC after: 2 weeks Sponsored by: BBOX.io
This fixes the build on other architectures where `struct x86_cpu_id` is unavailable. Sponsored by: The FreeBSD Foundation
Add the extended control interface Values match the reference ABI; the new struct layouts are pinned with static assertions. Reported by: thierry Reviewed by: adrian, thierry Differential Revision: https://reviews.freebsd.org/D59496
This is what FreeBSD now does for memdup_user and memdup_array_user (which is wrong for memdup, but correct for vmemdup). The amdgpu DRM driver started this in Linux ~6.12.86+ Reviewed by: wulf MFC after: 1 week Differential Revision: https://reviews.freebsd.org/D57444
[PATCH 06/31] FreeBSD OFED support for DPDK MLX5 PMD CQE compressing reduces PCI overhead by coalescing and compressing multiple CQEs into a single merged CQE. Successful compressing improves message rate especially for small packet traffic. CQE compressing is supported for all 64B CQE formats (with certain limitations) generated by RQ/Responder or by SQ/Requestor. IB/mlx5: Report mlx5 CQE compression capabilities when querying device The capabilities include: - Max number of compressed and aggregated CQEs in a single session, while zero means unsupported. - For Responder, there are two formats of mini CQE: mini CQE with Rx hash and mini CQE with checksum. They're mutual exclusive. Differential revision: https://reviews.freebsd.org/D32180 MFC after: 1 month
[PATCH 11/31] FreeBSD OFED support for DPDK MLX5 PMD The MLX5 core driver recognizes a VF device as a PF device in FreeBSD VM, because mlx5_core_pci_table was not having VF flag MLX5_PCI_DEV_IS_VF. Added MLX5_PCI_DEV_IS_VF flag in mlx5_core_pci_table. Updated PCI_VDEVICE macro in LinuxKPI for class and class_mask, same as linux. Differential revision: https://reviews.freebsd.org/D32185 MFC after: 1 month
Use designated initializers for .class and .class_mask in PCI_VDEVICE(), and set .driver_data explicitly on the VF entries, rather than relying on positional initializers. Sponsored by: NVidia networking MFC after: 1 month
PATCH 13/31] FreeBSD OFED support for DPDK MLX5 PMD The capabilities whether hardware support multi packet WQE or not is exposed to user space through query_device by uhw. Differential revision: https://reviews.freebsd.org/D32187 MFC after: 1 month
[PATCH 14/31] FreeBSD OFED support for DPDK MLX5 PMD a) Add support to receive specific Vxlan packet in ConnectX-4. b) Add support to match packet fields which are tunneled, i.e. support matching the header of the inner packet which is the result of or bit operation of the original header and the IB_FLOW_SPEC_INNER type. The combination of IB_FLOW_SPEC_INNER | IB_FLOW_SPEC_VXLAN_TUNNEL is not needed to be checked, because the IB core has this check already. Dfiferential revision: https://reviews.freebsd.org/D32188 MFC after: 1 month
[PATCH 15/31] FreeBSD OFED support for DPDK MLX5 PMD Enable mlx5 based hardware to report packet pacing capabilities from kernel to user space. Packet pacing allows to limit the rate to any number between the maximum and minimum, based on user settings. The capabilities are exposed to user space through query_device by uhw. The following capabilities are reported: - The maximum and minimum rate limit in kbps supported by packet pacing. - Bitmap showing which QP types are supported by packet pacing operation. Differential revision: https://reviews.freebsd.org/D32189 MFC after: 1 month
[PATCH 17/31] FreeBSD OFED support for DPDK MLX5 PMD Software parsing (SWP) is a feature that can be used to instruct the device to stop using its internal parser and to parse packets on the transmit path according to offsets set for each packets. Through this feature, the device allows the handling of checksum and LSO by the hardware according to the location of IP and TCP/UDP headers. Enable SW parsing on Raw Ethernet send queue by default if firmware supports it and report these capabilities to user space. Differential revision: https://reviews.freebsd.org/D32191 MFC after: 1 month
[PATCH 21/31] FreeBSD OFED support for DPDK MLX5 PMD a) This patch reports the device's striding RQ capabilities to the user-space: - min/max_single_stride_log_num_of_bytes: Log of min/max number of bytes in a single stride. - min/max_single_wqe_log_num_of_strides: Log of min/max number of strides in a single WQE. - supported_qpts: A bit mask to know which QP types support multi- packet RQ, for now only Raw Packet QPs. b) Allow creation of a multi-packet receive queue. In order to create a multi-packet RQ, the following fields in the mlx5_ib_rwq should be set: - log_num_strides: Log of number of strides per WQE - single_stride_log_num_of_bytes: Log of a single stride size - two_byte_shift_en: When enabled, hardware pads 2 bytes of zeros before writing the message to memory (e.g. for the IP alignment). Differential revision: https://reviews.freebsd.org/D32195 MFC after: 1 month
[PATCH 22/31] FreeBSD OFED support for DPDK MLX5 PMD a) The device can support receive Stateless Offloads for the inner packet's fields only when the packet is processed by TIR which is enabled to support tunneling. Otherwise, the device treats the packet as an ordinary non-tunneling packet and receive offloads can be done only for the outer packet's field. In order to enable receive Stateless Offloading support for incoming tunneling traffic the TIR should be created with tunneled_offload_en. Tunneling offloads is supported only be raw ethernet QP. This patch includes: - New QP creation flag for tunneling offloads. - Reports device capabilities. b) IB/mlx5: Add support for RSS on the inner packet Some user space application would like to do RSS on the inner packet fields instead on the outer. When MLX5_RX_HASH_INNER is set with one or more of the other hash fields, then the RSS will be done using the inner packet. Differential revision: https://reviews.freebsd.org/D32196 MFC after: 1 month
Import Linux upstream commit 4e2b53a5cb5a ("IB/mlx5: Report inner RSS
capability").
Define MLX5_RX_HASH_INNER as (1UL << 31). Shifting 1 into the sign bit of
a plain int is undefined behaviour, and user space compiles this header
too. The main.c half of that commit is already carried by the preceding
fixup.
Sponsored by: NVidia networking
MFC after: 1 month
[PATCH 25/31] FreeBSD OFED support for DPDK MLX5 PMD Struct mlx5_ib_striding_rq_caps was not aligned to 64 bit as it should have been. Add a 32 bit reserved field. Differential revision: https://reviews.freebsd.org/D32200 MFC after: 1 month
The security.mac.bsdextended.rules.<N> node handler takes N as `index = name[0]` (a signed int) and only checks `index >= MAC_BSDEXTENDED_MAXRULES`. A negative index is caught on the read branch, but the write-only add and delete branches proceed to `rules[index]` unconditionally. Reject `index < 0` alongside the existing upper-bound check. Submitted by calif.io for the OpenAI Patch The Planet program Signed-off-by: Andrew Griffiths <andrew@calif.io> Reviewed by: markj MFC after: 2 weeks
The kern.proc.kstack handler allocates its output and stack buffers before deliberately checking p_candebug again under the process lock. If trace-control state changes between authorization checks, the failure path balances the process and exec state but leaks both buffers. Submitted by calif.io for the OpenAI Patch The Planet program Signed-off-by: Andrew Griffiths <andrew@calif.io> Fixes: https://cgit.freebsd.org/src/commit/?id=8b5abd9027b8 ("kern_proc.c: disallow execve around sysctl kern.proc.kstacks") Reviewed by: markj MFC after: 1 week
Requred by drm-kmod v6.12.103 Obtained from: OpenBSD Reviewed by: bz, emaste MFC after: 1 week Differential Revision: https://reviews.freebsd.org/D59733
Required by drm-kmod v6.12.105 Reviewed by: bz, emaste MFC after: 1 week Differential Revision: https://reviews.freebsd.org/D59734
Add compile-time optional, non-sleeping fail points around every VF creation resource boundary, before VF VSI reconstruction, and in the GET_STATS validation path. Provide an ICE-wide wrapper and device selector so other driver subsystems can add scoped points without duplicating the failpoint plumbing. Keep the current SR-IOV points and VF selector in an iov child namespace. Compile the facility only with options DRIVER_FAILPOINTS. This shared option avoids a separate kernel option for every driver that provides test-only injection hooks. Ordinary kernels contain no ICE failpoint objects or sysctl nodes. Require an exact PF device name and optionally a VF index before any point can fire. This prevents a stale test setting from affecting another PF. The hooks exposed two reset-lifetime defects while validating the existing SR-IOV review series. A PF reset could discard a firmware VSI before teardown, and a rebuilt sibling could leave stale switch-filter state after IOV destroy. Tested on an E810-XXV with INVARIANTS and WITNESS. All twelve creation checkpoints rolled back and permitted immediate resource reuse. Forced reconstruction and malformed GET_STATS failures also preserved sibling operation and reply cardinality. Reviewed by: gallatin, nprice, ziaee MFC after: 2 weeks Sponsored by: BBOX.io Differential Revision: https://reviews.freebsd.org/D58940
This warning triggers in a few places in contributed code, and would therefore be annoying to fix. Use -Wno-error= to at least show the warnings so there is some incentive to submit them upstream. MFC after: 3 days
Now that semi-native drivers in main are clean of using any pre-4.15 Linux timer KPI, remove the macros from LinuxKPI as well. Supported versions of drm-kmod and the oldest nvidia port seem fine based on my checks. If anything else surfaces we can deal with the fallout; the changes to the newer KPI were mechanical in Linux and can be applied easily. This change can be detected by #if !defined(init_timer) at compile time so no need for a __FreeBSD_version bump. MFC after: 3 days Reviewed by: emaste Differential Revision: https://reviews.freebsd.org/D59688
Required by drm-kmod 6.12.101. Reviewed by: bz, dumbbell Differential Revision: https://reviews.freebsd.org/D59732
Reviewed by: olce Sponsored by: The FreeBSD Foundation Differential Revision: https://reviews.freebsd.org/D59739
This is intended to be used on resuming to check that everything that happened during suspend was expected. It is not meant to check for errors we should hard-fail from, rather it should be used as an opportunity for drivers such as amdsmu(4) to check e.g. whether the previous suspend-to-idle actually entered a deep sleep state (S0i3). Reviewed by: olce Sponsored by: The FreeBSD Foundation Differential Revision: https://reviews.freebsd.org/D59740
sched_4bsd: remove kern.sched.4bsd.followon kern.sched.4bsd.followon was enabled only for KSE. It was never compiled since ad1e7d285ab1 and KSE was removed years ago. Now it's time to remove this tunable. Reviewed by: olce Approved by: olce (mentor) MFC after: 2 weeks Sponsored by: FreeBSD Foundation Differential Revision: https://reviews.freebsd.org/D59394
sched_4bsd: remove dumping from maybe_preempt() 'dumping' is true only when kernel is dumping after crash (see minidumpsys()) so KERNEL_PANICKED() will catch this. Reviewed by: olce Approved by: olce (mentor) MFC after: 2 weeks Sponsored by: FreeBSD Foundation Differential Revision: https://reviews.freebsd.org/D59395
sched_4bsd: add static assertion for nice weight When NICE_WEIGHT * (PRIO_MAX - PRIO_MIN) exceeds the timeshare range, two CPU-bound threads with different nice values can have the same priority as their nice values are clamped to the timeshare range limit. Add static assertion on NICE_WEIGHT to ensure that there is always enough room for nice values in the both end of the timeshare priority range. Reviewed by: olce Approved by: olce (mentor) MFC after: 2 weeks Sponsored by: FreeBSD Foundation Differential Revision: https://reviews.freebsd.org/D59396
sched_4bsd: update function name in comment
In b43179fbe815 ("Create a new scheduler api..."), schedclock() was
renamed to sched_clock() but the function name in the comment remained
still. Update the comment to reflect up-to-date name for schedclock().
Reviewed by: olce
Approved by: olce (mentor)
Fixes: https://cgit.freebsd.org/src/commit/?id=b43179fbe815 ("Create a new scheduler api...")
MFC after: 2 weeks
Sponsored by: FreeBSD Foundation
Differential Revision: https://reviews.freebsd.org/D59397
sched_4bsd: remove obsolete comment In old Unix, the whole process address space including scheduler-related data was paged out to disk. We now allocate thread-related data with UMA on wired memory which never page out. Thus this comment is now obsolete. Reviewed by: olce Approved by: olce (mentor) MFC after: 2 weeks Sponsored by: FreeBSD Foundation Differential Revision: https://reviews.freebsd.org/D59398
Commit 62fa74d95a16 ("Add support for the new cpu...") removed all uses
of KTR_ULE, leaving the macro unused.
Reviewed by: olce
Approved by: olce (mentor)
Fixes: https://cgit.freebsd.org/src/commit/?id=62fa74d95a16 ("Add support for the new cpu...")
MFC after: 2 weeks
Sponsored by: FreeBSD Foundation
Differential Revision: https://reviews.freebsd.org/D59399
Fix three problems with kern.sched.{4bsd,ule}.slice:
* Guarantee minimum slice is 1.
* Recalculate hogticks on sysctl write.
* For ULE, recalculate sched_slice_min on sysctl write.
Reviewed by: olce
Approved by: olce (mentor)
MFC after: 2 weeks
Sponsored by: FreeBSD Foundation
Differential Revision: https://reviews.freebsd.org/D59401
sched_4bsd: fix comment in maybe_preempt() The comment says the new thread's priority is not a realtime priority while the code states pri > PRI_MAX_ITHD which is interrupt priorities not realtime. Reviewed by: olce Approved by: olce (mentor) MFC after: 2 weeks Sponsored by: FreeBSD Foundation Differential Revision: https://reviews.freebsd.org/D59402
sched_4bsd: move comment to correct location Reviewed by: olce Approved by: olce (mentor) MFC after: 2 weeks Sponsored by: FreeBSD Foundation Differential Revision: https://reviews.freebsd.org/D59403
sched_4bsd: remove obsolete comment 'awake' checks if a thread, not a process, is awake. Reviewed by: olce Approved by: olce (mentor) MFC after: 2 weeks Sponsored by: FreeBSD Foundation Differential Revision: https://reviews.freebsd.org/D59404
sched_4bsd: fix vague comment The comment "was incremented in schedcpu()" doesn't give enough background for decrementing ts_slptime by 1 (thus ignoring decay_cpu() for 1 ts_slptime). More accurately, ts_slptime is decremented by 1 because decay_cpu() has already executed once in schedcpu() when ts_slptime was 1. Reviewed by: olce Approved by: olce (mentor) MFC after: 2 weeks Sponsored by: FreeBSD Foundation Differential Revision: https://reviews.freebsd.org/D59406
In ULE ts_slice stores the number of ticks of slice passed not remaining. Reviewed by: olce Approved by: olce (mentor) MFC after: 2 weeks Sponsored by: FreeBSD Foundation Differential Revision: https://reviews.freebsd.org/D59407
sched_slice_min should always to be greater than zero. When modifying sched_slice through sysctl, if the new value is less than SCHED_SLICE_MIN_DIVISOR, sched_slice_min is computed to zero. Add imax(1, ...) to prevent this. tdq_slice() should not return a value less than sched_slice_min since that will cause integer underflow of ts2->ts_slice in sched_ule_fork_thread. SCHED_SLICE_MIN_DIVISOR is currently set to 6 so when load is 5 and sched_slice is 4, the two if conditions in tdq_slice() will pass and the function will return zero. Thus use imax() so tdq_slice returns sched_slice_min at minimum. Reviewed by: olce Approved by: olce (mentor) MFC after: 2 weeks Sponsored by: FreeBSD Foundation Differential Revision: https://reviews.freebsd.org/D59408
Suggested by: olce Reviewed by: olce Approved by: olce (mentor) MFC after: 2 weeks Sponsored by: FreeBSD Foundation Differential Revision: https://reviews.freebsd.org/D59471
sched_schedcpu() is called only during SYSINIT to start kthread that calls schedcpu() every second in 4BSD, but its name implies it's doing what 4BSD's schedcpu() does. Rename this function to sched_sysinit() to mark that schedulers can use it for its own SYSINIT routine. Note that their SYSINIT routine does not necessarily need to be similar to 4BSD's decay in schedcpu(). The scheduler.9 man page is planned to be rewritten from scratch, so no change to it for now. Reviewed by: olce Approved by: olce (mentor) MFC after: 2 weeks Sponsored by: FreeBSD Foundation Differential Revision: https://reviews.freebsd.org/D59437
To be used in _Generic() KPIs that accept different lock types. Reviewed by: kib Differential Revision: https://reviews.freebsd.org/D59455
This is type agnostic locked callout initializer. The callout_init_mtx() and etc remain for compatibility. Reviewed by: kib, markj Differential Revision: https://reviews.freebsd.org/D59456
And enable locking assertions for INVARIANTS kernel. Reviewed by: gallatin, kib, markj Differential Revision: https://reviews.freebsd.org/D59457
If a buffer ring experiences drops, then system is definitely starving in CPU resources. Having an extra cache miss on a racy shared variable increment doesn't help. Reviewed by: gallatin Differential Revision: https://reviews.freebsd.org/D59495
Without this patch, the code failed to acquire the error return for clnt_bck_rdma_send() and could use "error" uninitialized. No semantics change for non-RDMA NFS service. MFC after: 3 months Fixes: https://cgit.freebsd.org/src/commit/?id=7144a1d58c5c ("nfsd: Add glue for the nfsrdma.ko module")
Add a maxsegs field to rpcrdma_xprt, which is used to set the number of segments allowed for a chunk based on device attributes. This commit should not affect non-RDMA behaviour. MFC after: 3 months Fixes: https://cgit.freebsd.org/src/commit/?id=884ee8d6c9b4 ("nfscl: Add some glue for client side NFS over RDMA")
When compiling the kernel with gcc 14, errors similar to the following
are emitted:
sys/dev/cxgbe/iw_cxgbe/ev.c: In function 'c4iw_ev_handler':
sys/compat/linuxkpi/common/include/linux/xarray.h:132:23: error: statement with no effect [-Werror=unused-value]
132 | flags == 0; \
sys/dev/cxgbe/iw_cxgbe/ev.c:274:17: note: in expansion of macro 'xa_unlock_irqrestore'
274 | xa_unlock_irqrestore(&dev->cqs, flag);
| ^~~~~~~~~~~~~~~~~~~~
sys/compat/linuxkpi/common/include/linux/xarray.h:132:23: error: statement with no effect [-Werror=unused-value]
132 | flags == 0; \
sys/dev/cxgbe/iw_cxgbe/ev.c:283:17: note: in expansion of macro 'xa_unlock_irqrestore'
283 | xa_unlock_irqrestore(&dev->cqs, flag);
| ^~~~~~~~~~~~~~~~~~~~
It looks like the intent of the "flags == 0" statement was to make the
'flags' macro argument not unused, but it still results in a warning.
Use the construct "(void)flags" instead.
Reviewed by: bz
MFC after: 3 days
Differential Revision: https://reviews.freebsd.org/D59787
Calling dtrace_copy and dtrace_copystr with the kaddr and uaddr arguments inversed does not work with PAN. Rename them dtrace_copyin_pan and dtrace_copyinstr_pan, respectively, and implement dtrace_copyout_pan and dtrace_copyoutstr_pan. Avoid excessive faulting by checkin DTrace's CPU flags. Implement the trick from OpenSolaris/Illumos of only checking the flags when crossing into a new page, altough more effectively by examining the vaddr instead of the count. Reviewed by: markj MFC after: 3 weeks Differential Revision: https://reviews.freebsd.org/D59449
__func__ is a variable not a string literal so pass it to printf. This only manifest when KLD_DEBUG was defined so wasn't tested by an kernel including LINT. Reported by: Mark Millard <marklmi@yahoo.com> Sponsored by: Innovate UK
Without this, no kernel config contained this option so it was easy to break. Sponsored by: Innovate UK
Otherwise vfs_report_lockf() can race with lf_purgelocks() while the latter is freeing active lock entries without any locks held. Reviewed by: kib MFC after: 2 weeks Sponsored by: The FreeBSD Foundation Differential Revision: https://reviews.freebsd.org/D59768
This matches upstream behaviour since Linux's demangle_poll() silently discards POLLREMOVE from the requested events. Reviewed by: emaste Sponsored by: Sippy Software, Inc. Differential Revision: https://reviews.freebsd.org/D59914 MFC after: 1 week PR: https://bugs.freebsd.org/bugzilla/show_bug.cgi?id=297467
The Unicode closing single quotation mark is classified as a homoglyph and can trip automated code quality checks in downstream CI pipelines or cause code review UIs to refuse to display a file. If used as an apostrophe, use the ASCII single quote instead. If used as a closing single quote, replace with double quotes or no quotes at all. Sponsored by: Klara, Inc. Sponsored by: NetApp, Inc. Reviewed by: ziaee, obiwac, olce Differential Revision: https://reviews.freebsd.org/D59911
vfs_register hashes the filesystem name and uses it for sysctl oids. A filesystem name which hashes to 0 crashes in sysctl_register_oid(). Map 0 to 1 to prevent the kernel crash. This can be tested with "udf2" as the filesystem name. MFC after: 1 month MFC to: stable/15 stable/14 Reviewed by: kib Differential Revision: https://reviews.freebsd.org/D59839
If the entire plaintext is zero-filled, the backwards walk in tls13_find_record_type() would return the offset of the last byte of the TLS header. This causes an underflow when decrypting, resulting in a null pointer dereference. Fix the bug and add a regression test. Reviewed by: gallatin, jhb MFC after: 1 week Sponsored by: The FreeBSD Foundation Differential Revision: https://reviews.freebsd.org/D59767
For some reason, shm_partial_page_invalidate() unlocks the object upon an error, but its callers don't expect this. Don't do any special error handling. Keep the subroutine anyway since the name is a bit clearer than vm_page_grab_zero_partial(). While here, normalize the object pointer used for locking in shm_deallocate(). Reviewed by: kib Fixes: https://cgit.freebsd.org/src/commit/?id=454bc887f250 ("uipc_shm: Implements fspacectl(2) support") MFC after: 1 week Sponsored by: The FreeBSD Foundation Differential Revision: https://reviews.freebsd.org/D59877
put_device() already triggers lkpi_pci_dev_release(), which removes pdev from pci_devices, frees pdev->bus, destroys pcie_cap_lock, and uninits the DMA private data. Reported by: gallatin Fixes: https://cgit.freebsd.org/src/commit/?id=66b25ddf9125 ("LinuxKPI: pci detach: implement a proper detach (release) path") Reviewed by: kib, gallatin, bz Sponsored by: NVidia networking MFC after: 1 week Differential Revision: https://reviews.freebsd.org/D59913
All default jail parameter values are an empty or otherwise standard value, or are copied the jail's parent. A notable exception is the root directory, which is instead copied from the creating process's jail. Fix that to be in line with everything else. This change affects only the default when no path is specified; if a path of "/" is explicitly given, that will still be the creating process's root directory.
If we are inserting a run of pages into a VM object and fail at some point due to a memory allocation failure, we have to free all of the pages in the run. We do that by resetting some fields and calling vm_page_free_toq() on each page; this removes the page from the object and frees it back to the buddy allocator. If the page is supposed to be wired, we reset the reference count, but this was done incorrectly: the VPRC_OBJREF flag must be retained as the page still belongs to an object. Resetting it to zero will cause a panic in vm_page_free_prep(): vm_page_free_object_prep() will subtract VPRC_OBJREF from the refcount, causing underflow, and vm_page_free_prep() subsequently calls panic() if the refcount is non-zero. Reviewed by: alc, kib Fixes: https://cgit.freebsd.org/src/commit/?id=fee2a2fa3983 ("Change synchonization rules for vm_page reference counting.") MFC after: 1 week Sponsored by: The FreeBSD Foundation Differential Revision: https://reviews.freebsd.org/D59908
A race is possible otherwise: vfs_busy() may return after an unmounted filesystem has been removed from the global mount list. That is, vfs_busy() will block until vfs_mount_destroy() sets MNTK_REFEXPIRE, and at that point the mountpoint has been removed from the mountlist, so TAILQ_FOREACH can return an invalid value. Simply do not block if the mountpoint is being unmounted. Reviewed by: kib Fixes: https://cgit.freebsd.org/src/commit/?id=eca39864f702 ("Add sysctl KERN_LOCKF") MFC after: 1 week Sponsored by: The FreeBSD Foundation Differential Revision: https://reviews.freebsd.org/D59982
that locks both bo vnode and topology. Reviewed by: markj Tested by: pho Sponsored by: The FreeBSD Foundation MFC after: 1 week Differential revision: https://reviews.freebsd.org/D59932
Stop exempting devfs nodes. Reviewed by: markj Tested by: pho Sponsored by: The FreeBSD Foundation MFC after: 1 week Differential revision: https://reviews.freebsd.org/D59932
It does not provide value over the pre-existing SYSCTL_STATIC_CHILDREN() macro, which is widely used in the tree, whereas SYSCTL_NODE_CHILDREN() is not. Reviewed by: markj MFC after: 3 days Sponsored by: The FreeBSD Foundation Differential Revision: https://reviews.freebsd.org/D60009
This makes it easier to have a full list of them at a glance, and makes for a natural place to add new ones. MFC after: 3 days Sponsored by: The FreeBSD Foundation Differential Revision: https://reviews.freebsd.org/D60010
NOTES: SMP: Move PREEMPTION out of the debugging options section It has been activated by default for more than 20 years. Reviewed by: mchoo, srcmgr (imp) Fixes: https://cgit.freebsd.org/src/commit/?id=444ba945136b ("Switch the default scheduler to 4BSD...") MFC after: 3 days Sponsored by: The FreeBSD Foundation Differential Revision: https://reviews.freebsd.org/D60012
NOTES: IPI_PREEMPTION: Fix documentation, applies to all architectures Move its description into 'sys/conf/NOTES' and update it to match reality. Reviewed by: scheduler (mchoo) MFC after: 3 days Sponsored by: The FreeBSD Foundation Differential Revision: https://reviews.freebsd.org/D60013
NOTES: Multiple schedulers can be compiled in at once Mention that ULE is the default scheduler when multiple ones are compiled in and how the tunable 'kern.sched.name' can be used to select another one. While here, regroup SCHED_ULE and SCHED_4BSD, as they control if the respective scheduler instances are compiled in, putting SCHED_STATS aside. Reviewed by: mchoo Fixes: https://cgit.freebsd.org/src/commit/?id=75a66a92c92f ("- Add an option to compile in SCHED_STATS. ...") Fixes: https://cgit.freebsd.org/src/commit/?id=1322760fd127 ("sys: enable both SCHED_ULE and SCHED_4BSD for some configs") MFC after: 3 days Sponsored by: The FreeBSD Foundation Differential Revision: https://reviews.freebsd.org/D60014
The corresponding machinery was removed in commit 0c4440c3aafe6, ten yearso ago. No functional change intended. Reviewed by: imp Sponsored by: The FreeBSD Foundation Differential Revision: https://reviews.freebsd.org/D59986
- refcount_acquire() returns the old value, use that to assert that the old value was non-zero. - refcount_release() already asserts that the refcount value is non-zero, so don't bother asserting that again in the jail code. - Use __diagused instead of having separate implementations for INVARIANTS and !INVARIANTS. No functional change intended. Reviewed by: jamie MFC after: 1 week Sponsored by: The FreeBSD Foundation Differential Revision: https://reviews.freebsd.org/D59983
If we're killing a jail which has some user refs pending, then we would first drop our ref and then kill all processes in the prison. However, it's possible for the prison to be freed before we finish that operation, generally if the processes exit on their own before prison_proc_iterate() returns. Thus, defer the release of the prison refcount until after we've killed all procs. Reviewed by: jamie MFC after: 2 weeks Sponsored by: The FreeBSD Foundation Differential Revision: https://reviews.freebsd.org/D59984
atomic_load_* is a read, not a write. Otherwise KASAN will report a use-after-free via atomic_load_* as a write rather than a read. MFC after: 1 week Sponsored by: The FreeBSD Foundation
aout_coredump() was removed in commit 1eecfae3e53cb3. No functional change intended. MFC after: 1 week Sponsored by: The FreeBSD Foundation
nlmsg_translate_ifname_nla() always used nw->ifp for the name, which is fine for a single-message ifnet event, but an RTM_GETLINK dump holds one RTM_NEWLINK per interface and is translated with the ifp the writer had when the buffer was flushed. The root of the problem is that msgs_to_linux() takes a single ifp for a buffer that may contain messages about many interfaces. Use ifi_index to resolve the name instead. Reviewed by: glebius Differential Revision: https://reviews.freebsd.org/D59595
We were not acquiring the global sysvshm lock when handling cleanup of sysvshm segments. Acquire the lock in shm_prison_cleanup() instead, to be consistent with the sysv semaphore code. Reviewed by: jamie MFC after: 1 week Sponsored by: The FreeBSD Foundation Differential Revision: https://reviews.freebsd.org/D60030
To allow normal users to render through the render device, we expose the renderD node. This enables Wayland applications to use hardware acceleration when running under the Linux emulator. Additionally, libdrm and Mesa need to look up /sys/dev/char/<major>:<minor> and <pcidev>/drm to identify the corresponding renderer device (e.g., a renderD device). We expose this path as well so that libdrm can locate the renderer. Differential Revision: https://reviews.freebsd.org/D59190
This aligns with copyinuio_t and avoids some hypothetical undefined behavor around calling functions with mismatched types. Reviewed by: kib Effort: CHERI upstreaming Sponsored by: Innovate UK Differential Revision: https://reviews.freebsd.org/D60024
At present, PGA_EXECUTABLE is only used by the powerpc mmu_oea64 pmap. The MI layer only accesses this flag to assert that it is clear when a managed page is freed. Soon, we will need a similar, but not identical, machine-dependent flag in the arm64 pmap. So, we rename this flag to PGA_PMAP_PRIV1, simply saying that it is reserved for use by the pmap. Each pmap can then define a name that best reflects its own meaning. However, we still assert that this flag is clear when a managed page is freed. No functional change. Reviewed by: kib, markj Differential Revision: https://reviews.freebsd.org/D59995
Since 80c7315d17ce ("Restore signal mask in epoll_pwait.") the caller's
signal mask is saved in a local variable, but TDP_OLDMASK is still set
and the TDA_SIGSUSPEND AST is still scheduled. On return to user mode
that AST, or postsig() if a signal is delivered, then installs
td_oldsigmask, which this code never writes and which holds whatever
mask the thread had at its last sigsuspend(2), pselect(2) or ppoll(2).
As a result every epoll_pwait(2)/epoll_pwait2(2) call with a non-NULL
sigmask can leave the thread with a stale signal mask. In addition, the
explicit restore at the end overwrote the return value, so EINTR (and
any other error) was reported to user space as 0.
Save the old mask in td_oldsigmask and let the AST restore it, as
kern_pselect() and kern_poll_kfds() do: schedule TDA_SIGSUSPEND if the
wait was interrupted, so the signal is delivered with the temporary mask
in place, and TDA_PSELECT otherwise. This matches Linux, which restores
the saved mask unless the syscall returns -EINTR.
This deadlocks Bun-based programs such as Claude Code (>= 2.1.269) and
opencode. JavaScriptCore suspends threads for conservative GC stack
scanning by sending SIGPWR and waiting for the target's handler, which
calls sigsuspend(2) with SIGPWR blocked. Afterwards td_oldsigmask
contains SIGPWR, the event loop's next epoll_pwait2(2) (Bun always
passes an empty sigmask) blocks SIGPWR on the main thread, and the next
GC suspend request waits forever.
PR: https://bugs.freebsd.org/bugzilla/show_bug.cgi?id=298878
Fixes: https://cgit.freebsd.org/src/commit/?id=80c7315d17ce ("Restore signal mask in epoll_pwait.")
MFC after: 1 week
This keeps it in sync with sys/modules/ibcore/Makefile and fixes the build on GCC 16 after the changes to -Wunused*[0]. [0] https://gcc.gnu.org/gcc-16/porting_to.html#changes-to-wunused Reviewed by: jhb MFC after: 3 days Sponsored by: The FreeBSD Foundation Differential Revision: https://reviews.freebsd.org/D59539
We have always returned EPERM in this case, but it's incorrect, we should return ECAPMODE for capability mode violations. Fix the errno value. Reviewed by: emaste MFC after: 2 weeks Differential Revision: https://reviews.freebsd.org/D59887
The subsequent namei() call is relative to AT_FDCWD, and such lookups are always disallowed in capability mode. No functional change intended. Reviewed by: emaste Differential Revision: https://reviews.freebsd.org/D59888
Forcing maxsize to RPC_MAXDATASIZE breaks code that deliberately uses XDR with longer strings than permitted by SunRPC. MFC after: 1 week Fixes: https://cgit.freebsd.org/src/commit/?id=6448ec89e739 (" * limit size of buffers to RPC_MAXDATASIZE") Fixes: https://cgit.freebsd.org/src/commit/?id=e17d7ab869bb ("xdr_string: don't leak strings with xdr_free") Sponsored by: Klara, Inc. Sponsored by: NetApp, Inc. Reviewed by: kevans, brooks Differential Revision: https://reviews.freebsd.org/D59994
semop() may sleep waiting for a semaphore. Upon waking up, it checks to see if the set's sequence number has changed, indicating that the set was removed. The sequence number is not wide enough to prevent a false negative due to wraparound, in which case the subsequent access of `semakptr->u.__sem_base[sopptr->sem_num]` may be out of bounds. This race can be leveraged to elevate privileges. Fix this by introducing a 64-bit sequence number for each semaphore pool. This is wide enough to make the race impossible to hit. Allocate a separate array for them, as we cannot really change the layout of struct semid_kernel since some userspace tools (e.g., ipcrm(1)) embed the layout. While here, use semvalid() instead of open-coding its implementation, convert a couple of flags to be bool, and use a better variable name to store required permissions. Approved by: so Security: FreeBSD-SA-26:64.sysvsem Security: CVE-2026-58098 Reported by: Reo Shiseki Reported by: Andrew Griffiths Reviewed by: kib Sponsored by: The FreeBSD Foundation Differential Revision: https://reviews.freebsd.org/D59347
Commit d8bdcb08d0eb fixed a problem in kqueue_fork_copy_knote() where we did not skip over marker knotes when copying. However, that fix was not sufficient: we bump the influx counter and check for a marker after dropping the kqueue lock. So, if multiple threads in a process are forking concurrently, kqueue_fork_copy_list() may mark a marker as in-flux and drop the lock; if the marker owner then frees the marker, the first thread will decrement the in-flux counter of a freed knotes. This use-after-free can be exploited, at least prior to commit d8bdcb08d0eb, which makes exploitation more challenging. Approved by: so Security: FreeBSD-SA-26:65.kqueue Security: CVE-2026-58099 Reported by: Reo Shiseki Reviewed by: kib Sponsored by: The FreeBSD Foundation Differential Revision: https://reviews.freebsd.org/D59522
Here, fdp points to the new fdtable, copied from that of the parent process. There is a window after the fdtable is copied, and before kqueue_fork_copy_knote() runs, where a different thread in the parent could have grown the parent's fdtable and registered a knote with ident larger than the size of the child's fdtable. This race can lead to an out-of-bounds read. Add a bounds check for this case; skip the knote if it is referencing a non-existent file. Approved by: so Security: FreeBSD-SA-26:65.kqueue Security: CVE-2026-58100 Reviewed by: kib Sponsored by: The FreeBSD Foundation Differential Revision: https://reviews.freebsd.org/D59916
In a couple of places we want to know whether someone has limited rights on an fd. There, we want a predicate which determines whether the set of rights is smaller than CAP_ALL, and whether there are explicit ioctl or fcntl lists. Factor this out into a helper function, in preparation for use elsewhere. No functional change intended. Approved by: so Security: FreeBSD-SA-26:66.jail Reviewed by: kib Sponsored by: The FreeBSD Foundation Differential Revision: https://reviews.freebsd.org/D59884
These routines let one compute the intersection of two sets of filecaps or capability rights, just as filecaps_merge() and cap_rights_merge() compute the union. This will be useful in an upcoming patch. filecaps_intersect() is complex due to the need to merge sets of ioctls. For now this is implemented with a dumb nested loop on the basis that ioctl lists are typically short enough that this is fine. It may be better to instead sort the two lists first and step through them together. No functional change intended. Approved by: so Security: FreeBSD-SA-26:66.jail Reviewed by: kib Sponsored by: The FreeBSD Foundation Differential Revision: https://reviews.freebsd.org/D59885
The FD_RESOLVE_BENEATH flag was intended to try to resolve bugzilla PR 262179 without entirely disallowing fd passing between jails. However, one can use renameat() to bypass the restriction: upon receiving a directory fd with FD_RESOLVE_BENEATH set, a jailed process can still move its CWD or one of its ancestors to the directory, and just cd out of its jail root. So disallow renameat() when either the source or destination directory fds has FD_RESOLVE_BENEATH set, like we do with fchdir() and fchroot() to prevent similar escapes. Approved by: so Security: FreeBSD-SA-26:66.jail Security: CVE-2026-101305 PR: https://bugs.freebsd.org/bugzilla/show_bug.cgi?id=262179 Reported by: firk@cantconnect.ru Reviewed by: olce, kib Differential Revision: https://reviews.freebsd.org/D59875
This function take a struct uio previously created by copyinuio and and updates the lengths of the user-space iovec to match those in the uio. To reduce the risks of pointer leakage and cross-ABI pointer confusion, lengths are updated individually. Reviewed by: jamie, jhb Effort: CHERI upstreaming Sponsored by: DARPA, AFRL, Innovate UK Differential Revision: https://reviews.freebsd.org/D60025
This will enable compat implementations in the future. Reviewed by: jamie, jhb Effort: CHERI upstreaming Sponsored by: Innovate UK Differential Revision: https://reviews.freebsd.org/D60026
I suspect this overflow check isn't needed at all, at least today, since vm_map_insert() will detect wraparound when it creates segments (and even a check against UINT_MAX is too loose, since the max user address is AOUT32_USRSTACK == 0xbfc00000). But if we're going to keep this check, there doesn't seem to be any downside to applying it on all platforms, before exec_new_vmspace() tears down the current vmspace. Reported by: Muhammed Sariyildiz <asiyee994@gmail.com> Reviewed by: kib MFC after: 2 weeks Sponsored by: The FreeBSD Foundation Differential Revision: https://reviews.freebsd.org/D60027
vm_swapout: Restore handling of RLIMIT_RSS In commit 13a1129d700c, I removed the mechanism by which the pagedaemon signals the swapout thread when page reclamation is unable to keep up with demand. This is because the swapout thread's main action in this case is to swap out sleeping processes, but we removed this support. However, it had the secondary effect of causing the swapout thread to enforce RLIMIT_RSS when racct is not enabled. Without it, if racct_enabled is false, nothing ever kicks the swapout thread. Restore the old behaviour of trying to enforce RLIMIT_RSS when the page daemon is unable to keep up with demand. I'm not at all convinced this is a good way to implement the limit, but the change wasn't intentional, so let's restore it for now. Fixes: https://cgit.freebsd.org/src/commit/?id=13a1129d700c ("vm: Remove kernel stack swapping support, part 1") Reviewed by: olce, kib MFC after: 2 weeks Differential Revision: https://reviews.freebsd.org/D56142
vm_swapout: Fix the build without options RACCT Reported by: Jenkins Fixes: https://cgit.freebsd.org/src/commit/?id=da0764f23554 ("vm_swapout: Restore handling of RLIMIT_RSS")
amd64: Revert an unintended change to GENERIC Fixes: https://cgit.freebsd.org/src/commit/?id=767cd9bb4200 ("vm_swapout: Fix the build without options RACCT")
The freebsd14_setgroups() function would try to modify the effective GID on the current process' credentials without holding the process lock, allowing races with other threads concurrently modifying the process credentials. In the worst case, freebsd14_setgroups() could be manipulating a 'struct ucred' already freed by another thread (in the very small window after reading 'p_ucred' without lock but before modifying the effective GID). Concurrent uses of freebsd14_setgroups() or setcred() could also lead to non-atomic credentials modifications. Fix this by making kern_setgroups() take a new boolean indicating whether the passed array includes the effective GID in its first slot. When this boolean is true, it internally keeps the effective GID in a separate variable, pretends that the groups[] array that was passed actually starts at 'groups + 1', do the usual steps to set the supplementary groups and new extra ones to set the effective GID along, without releasing the process lock in between. Reported by: markj Reviewed by: markj MFC after: 2 weeks Sponsored by: The FreeBSD Foundation Differential Revision: https://reviews.freebsd.org/D60028
MFC with: e6fbef451dd4 ("cred: Fix a race in the FreeBSD-14-compatible setgroups(2)")
Sponsored by: The FreeBSD Foundation
epoch_trace_report() assigned the return value of RB_INSERT() back to the new element. When two threads report the same stack concurrently, the loser's RB_INSERT() returns the element already in the tree, and that element was freed while still linked, leaking the new allocation. The next lookup touches freed memory; KASAN catches it as a use-after-free. Keep the return value separate and free the new element instead. The thread that won the race prints the report, so return without printing it a second time. Reviewed by: markj Fixes: https://cgit.freebsd.org/src/commit/?id=173c062a569b ("Improve EPOCH_TRACE") MFC after: 1 week Sponsored by: Rubicon Communications, LLC ("Netgate") Differential Revision: https://reviews.freebsd.org/D60162
This is consistent with other system call helpers. Reviewed by: markj Sponsored by: AFRL, DARPA Differential Revision: https://reviews.freebsd.org/D60173
Reviewed by: jah, sobomax, markj, olce Sponsored by: The FreeBSD Foundation MFC after: 1 week Differential revision: https://reviews.freebsd.org/D59910
This update brings spdxtool(1), with the ability to generate software bill of material files (SBOM) in the SPDX 3.0.1 format (JSON-LD). Reviewed by: markj Approved by: markj Sponsored by: The FreeBSD Foundation Differential Revision: https://reviews.freebsd.org/D57953
We have separate ports for Ccache 3 and 4. Suggest both, rather than only the Ccache 3 port. Rearrange the text somewhat to avoid an excessively ragged edge on a standard 80-column terminal. MFC after: 1 week Reviewed by: markj Differential Revision: https://reviews.freebsd.org/D58005
On most architectures we end up not needing ABIBreak.cpp as, although some of the sources here do reference EnableABIBreakingChecks (or, if assertions are disabled, DisableABIBreakingChecks) at a source level, we compile with -ffunction-sections and -fdata-sections, and link with --gc-sections, and it happens to be the case that all references can be GC'ed. However, prior to LLVM 21, the RISC-V backend did not apply -fdata-sections to .sdata, where references to these symbols end up, and for some files we're building with such references we end up not being able to GC .sdata due to the other unrelated data in it, meaning that we do in fact need to build ABIBreak.cpp. Whilst we could make this conditional on the architecture, it's a tiny file, and it's a bit fragile to rely on GC behaviour, so just include it unconditionally. Reviewed by: dim, emaste Fixes: https://cgit.freebsd.org/src/commit/?id=770cf0a5f02d ("Fixups after llvm-project main llvmorg-21-init-19288-gface93e724f4 merge") MFC after: 1 week Differential Revision: https://reviews.freebsd.org/D58044
Reviewed by: markj Sponsored by: The FreeBSD Foundation MFC after: 1 week Differential revision: https://reviews.freebsd.org/D57124
Retire the GNU subtree With GNU diff and cdialog gone, this is now an empty shell. Reviewed by: markj Differential Revision: https://reviews.freebsd.org/D55425
Add a few missed files to ObsoleteFiles.inc There were still some left-over files under usr/tests/gnu/usr.bin/diff, causing the directory to not be fully removed. Add these to OLD_FILES. Fixes: https://cgit.freebsd.org/src/commit/?id=134a4c78d070
Reviewed by: markj Tested by: pho Sponsored by: The FreeBSD Foundation MFC after: 1 week Differential revision: https://reviews.freebsd.org/D57163
-I${SRCTOP}/sys/contrib/xz-embedded/linux/lib/xz isn't used, and
.PATH: ${SRCTOP}/sys/contrib/xz-embedded/freebsd isn't used either.
Remove them both to simplify things a little.
Sponsored by: Netflix
The new world may use system calls that are not in the currently-running
kernel, so we cannot chroot into the new environment to run `make
installworld`, `etcupdate`, etc. Partially revert commit 16702050ac95
("beinstall: perform pre-installworld steps") and switch back to using
DESTDIR for installworld and so on.
Reported by: olivier
Reviewed by: olivier
Sponsored by: The FreeBSD Foundation
Differential Revision: https://reviews.freebsd.org/D50682
I introduced it in commit 1b49115a40ad ("Promote llvm-cov to a
standalone option"). llvm-cov was previously enabled as part of the
CLANG_EXTRAS option. I made it a standalone, default-enabled option for
parity with the tools provided by the GCC-based toolchain.
We no longer provide an in-tree GCC toolchain. Now, just build llvm-cov
along with Clang to simplify build infrastructure.
Reviewed by: dim
Sponsored by: The FreeBSD Foundation
Differential Revision: https://reviews.freebsd.org/D58155
RANLIB is not used by our build, so there is no need to set it. Reviewed by: imp Sponsored by: The FreeBSD Foundation Differential Revision: https://reviews.freebsd.org/D58156
For the BIOS, add xzfs support. This is a tiny increase in the loader size, but allows us to fetch compressed files from any of the filesystems we support, including over the network. For EFI, also add gzipfs and bzip2fs support we well. The increment for these files is tiny. Sponsored by: Netflix
Sponsored by: Netflix
The tcp_bblog facility provides structured logging of TCP stack activity for debugging and performance analysis. It is implemented in the kernel and allows per-connection tracing of TCP events with low overhead. Reviewed by: tuexen, ziaee MFC after: 1 week Relnotes: yes Differential Revision: https://reviews.freebsd.org/D56252
sbintime.9 is a manual page that documents the usage of sbintime_t and its helper functions. MFC after: 1 week Reviewed by: ziaee, markj Differential Revision: https://reviews.freebsd.org/D57931
The debugfs options between the various modules (core and chipsets) are not 100% de-coupled. This means we may run into unresolveable symbols at load time of the modules if we enable certain options generally or for core but not for the chipset. For now: always build the core module with debugfs support. Migrate the CONFIG_MAC80211_DEBUGFS flag into the Makefile of each chipset so we can individually turn it on. Sponsored by: The FreeBSD Foundation MFC after: 3 days
This is included via acpivar.h so needs to be in SRCS to be generated. Reported by: bz Fixes: https://cgit.freebsd.org/src/commit/?id=bc49842769bd ("acpi_einj: Support for ACPI error injection") Sponsored by: Arm Ltd
Add a workaround for the Arm Cortex-A53 erratum 843419. This has been targeted when the build is either unoptimised for any CPU/architecture or targets the Cortex-A53 or ARMv8.0 architecture. PR: https://bugs.freebsd.org/bugzilla/show_bug.cgi?id=296240 PR: https://bugs.freebsd.org/bugzilla/show_bug.cgi?id=296395 Reported by: Hal Murray <halmurray+freebsd@sonic.net> Reported by: Andreas Schuh <x55839@icloud.com> Reviewed by: cognet, mmel Sponsored by: Arm Ltd Differential Revision: https://reviews.freebsd.org/D58212
Since the devd rules use sysrc, bsdconfig should be installed. MFC after: 3 days
Until D57524 is not reviewed and committed we will have a missing function declaration which prevents us to compile (in) debugfs for mt76 core and mt7921. Temporary disable debugfs again. Sponsored by: The FreeBSD Foundation MFC after: 3 days
We defined CONFIG_DEBUG_FILE only in libwpautils, not in wpa_supplicant, so all it did was enable code that never got called. Enable it at the top level so it also applies to wpa_supplicant(8), and the -f option mentioned in the manual page now actually works. PR: https://bugs.freebsd.org/bugzilla/show_bug.cgi?id=281617 MFC after: 1 week Reviewed by: cy Differential Revision: https://reviews.freebsd.org/D57723
The ssh-sk-helper utility only functions if/when MK_USB == yes. Installing it on systems where MK_USB == no doesn't make sense. Differential Revision: https://reviews.freebsd.org/D58246
Commit 1876de606eb8 exposed missing symbols that the port security/krb5
installed that the base system did not install. Part of the solution
was to make libprofile.so private (not libprofile.a) just as the port
does, Red Hat Enterprise Linux does, and as installing MIT KRB5 by hand
does. The actual fix for this was to put symbols and their corresponding
functions into the correct librarires, i.e. libkrb5.so and othes, just
as the port, Red Hat, and manually installed via tarball do.
Unfortunately INTERNALLIB disables the include of bsd.incs.mk and the
install of header files. This is still needed to install profile.h into
/usr/include (just as the port installs it into ${LOCALBASE}/include
and RHEL installs it in /usr/include). This commit fixes this by
installing profile.h into /usr/include from the krb5/include Makfile.
Reported by: fluffy
Tested by: fluffy
Reviewed by: fluffy
Fixes: https://cgit.freebsd.org/src/commit/?id=1876de606eb8
MFC after: 3 days
Differential Revision: https://reviews.freebsd.org/D58286
Reviewed by: markj Sponsored by: The FreeBSD Foundation MFC after: 1 week Differential revision: https://reviews.freebsd.org/D58463
We used to pass CONFIGURE_ARGS to the make command which builds pkg,
but ports/ports-mgmt/pkg/Makefile has its own CONFIGURE_ARGS and the
version we were providing at the command line didn't contain the
--mandir setting which was added to the port with pkg 2.8.0. This
broke release builds.
Instead of passing --prefix=${LOCALBASE} via CONFIGURE_ARGS, pass
PREFIX=${LOCALBASE}; the port Makefile passes that value through to
its configure script. We also used to pass a --host parameter, but
that seems to have become unnecessary at some point in the past decade.
MFC after: 1 day
Sponsored by: Amazon
This module has several source files, with many conditional on the platform architecture. Make it easier to read, and better for future diffs against these lists. - Convert to one SRC per line - Simplify arm/armv7 condition - Remove now-empty header comment - Minor formatting MFC after: 1 week Sponsored by: The FreeBSD Foundation
Reviewed by: markj Differential Revision: https://reviews.freebsd.org/D58531
Document supported controllers, PF and VF naming, PCI_IOV and IOMMU requirements, queue and lifecycle constraints, iovctl schema, filtering and anti-spoof policy, mailbox and MDD recovery, shared hardware limits, rate control, and statistics cadence. Relnotes: yes Sponsored by: BBOX.io
Google Cloud recommends migrating from gsutil to gcloud storage CLI. Update gce-do-upload target to use `gcloud storage buckets create` and `gcloud storage cp` instead of `gsutil mb` and `gsutil cp` commands. PR: conf/https://bugs.freebsd.org/bugzilla/show_bug.cgi?id=297016 Reviewed by: lwhsu MFC after: 3 days Differential Revision: https://reviews.freebsd.org/D58464
If one of the source files we copy is non-writeable, cp will create a non-writeable copy. If the original is later modified, cp will fail to overwrite the copy since it is not writeable. Using cp -f ensures the copy always succeeds, as long as the object directory is writeable. MFC after: 1 week Sponsored by: Klara, Inc. Sponsored by: NetApp, Inc.
Increase EFI partition size to begin rootfs at 64mb. I believe this was my original intention. I have a microSD card with 8mb block size which emits an advisory in verbose dmesg about the misaligned partition. MFC after: 1 week Sponsored by: The FreeBSD Foundation
uart(4), unix(4), veriexec(4), video(4) and the gzero(4) MLINK are
not USB things, but they were in the .if ${MK_USB} != "no" block.
So if we build with WITHOUT_USB, these man pages are lost. Move
them out of the block.
MFC after: 3 days
Sponsored by: The FreeBSD Foundation
Document supported virtual-function families, driver features, queue negotiation, PF-controlled policy, and media limitations. MFC after: 2 weeks Sponsored by: BBOX.io
Parts of rtentry.9 information such as information related to the nexthop is outdated. Remove those relics and add the new design into separate manual instead. Reviewed by: bcr Discussed with: ziaee Differential Revision: https://reviews.freebsd.org/D58564
Reported by: Herbert J. Skuhra
The module's i386 source list compiles the files that call pmc_rapl_initialize() and pmc_rapl_finalize() but not the one that defines them, so the i386 hwpmc.ko has both undefined and cannot be loaded. Fixes: https://cgit.freebsd.org/src/commit/?id=a99d04f39dab ("hwpmc: add RAPL energy-counter class (AMD + Intel)") Assisted-by: Claude Code (Opus 5)
Enable the new growfs_postboot mechanism. Note that this also implies disabling automatic allocation of swap space on the root disk, since we cannot grow the root filesystem if swap space is allocated after it. This will not be MFCed since it is a significant behavioural change. Sponsored by: Amazon Relnotes: yes
A recent commit started to use _uid and _gid in <bsd.dirs.mk> to mangle the user and group for newly installed directories when MK_INSTALL_AS_USER is set. However, _gid was previously only set when _uid was not 0, causing the group to be set to the empty string. Set _uid and _gid together to avoid this problem. Fixes: https://cgit.freebsd.org/src/commit/?id=541e6e2d516b6c9d3681b24464e9ef53c1f2579a PR: https://bugs.freebsd.org/bugzilla/show_bug.cgi?id=297841 Reviewed by: emaste, imp Reported by: Ralph Zitz <ralph@zitz.dk> Differential Revision: https://reviews.freebsd.org/D59150
This mirrors the changes applied to kernel debug symbol sets in commit 9a354a41be9a40c3c0a16cc20f4009d3b31679cc. Reviewed by: emaste Sponsored by: AFRL, DARPA Differential Revision: https://reviews.freebsd.org/D59058
MFC after: 1 month MFC to: stable/15 MFC to: stable/14
Sponsored by: Netflix
- Remove trailing whitespace - Fix typo with variable referenced adding sources for `test_fuzz`. MFC after: 2 weeks
This test has not passed since 185becb1e1bd2657c156f78aeb52edac05ba5fb5 (the libarchive 3.8.9 upgrade). PR: https://bugs.freebsd.org/bugzilla/show_bug.cgi?id=273732 MFC after: 1 week Reviewed by: siva Differential Revision: https://reviews.freebsd.org/D59314
In 16.0/15.1, the PAM modules were split from FreeBSD-runtime into a new FreeBSD-pam package. FreeBSD-runtime does not install FreeBSD-pam, which means if a user starts from runtime, then installs sshd, sshd will fail to authenticate users because of missing PAM modules. Since FreeBSD-pam is relatively small (about 230kB on amd64), and is already part of FreeBSD-set-minimal, add it to the runtime image as well. Users who absolutely don't want this can still build their own images without it. MFC after: 1 week Reviewed by: dfr Reported by: Michael Johnson <ahze@ahze.net> Sponsored by: The FreeBSD Foundation Differential Revision: https://reviews.freebsd.org/D59194
By default, makefs uses the host environment's user and group databases
when creating filesystems. This causes makefs to fail when trying to
create files owned by users or groups which don't exist in the host
environment, for example when creating a VM with packages pre-installed
which added their own users/groups.
Pass "-N ${DESTDIR}/etc" to makefs to point it at the user and group
databases from the image being created.
MFC after: 1 week
Sponsored by: Amazon
That function was removed in commit 98549e2dc6fb0. PR: https://bugs.freebsd.org/bugzilla/show_bug.cgi?id=288081 Fixes: https://cgit.freebsd.org/src/commit/?id=98549e2dc6fb ("Centralize the logic in vfs_vmio_unwire() and sendfile_free_page().")
ATF_TESTS_C automatically adds the appropriate library to LDADD -- there's no need to manually append the same library. MFC after: 1 week Differential Revision: https://reviews.freebsd.org/D59411
Desktop AMIs have xrdp enabled and boot to a KDE desktop; they are as compatible as possible with EC2 Windows AMIs, setting a random password and printing it to the console in encrypted format to be retrieved using the EC2 GetPasswordData API. Two rc.d scripts are included in this commit which will not exist in the long term: ec2_addpass will become part of the ec2-scripts package, and ec2_desktop_extras will go away once its functionality is included elsewhere. MFC After: 1 month Relnotes: yes Sponsored by: Amazon
which makes the variable available for machine/Makefile.inc usage. Reviewed by: fuz, kfv Sponsored by: The FreeBSD Foundation MFC after: 1 week Differential revision: https://reviews.freebsd.org/D59462
Build with 1.5 * ncpu jobs instead of hardcoding 40. Reviewed by: imp Differential Revision: https://reviews.freebsd.org/D59515
BEARSSL is disabled by default, add this here to make sure it doesn't break. Reviewed by: imp Differential Revision: https://reviews.freebsd.org/D59516
Give the WITHOUT_LOADER_BIOS_TEXTONLY build its own log file.
This module is intended to compile in its own directory. Therefore, the SYSDIR is not defined by default. See 0433870efefc. Fixes: https://cgit.freebsd.org/src/commit/?id=8e985774117d ("kern: Remove needless kern.opts.mk")
echo is a shell builtin, and always present. While ls usually works, some upgrade scenarios need to have ls be a build to to work correctly. Rather than add it as a build tool, use echo instead which is always safe because sh is a build tool. Sponsored by: Netflix
For consistency, create a symbolic link from mlx5en.4 to also if_mce.4 Reviewed by: ziaee, #manpages Event: EuroBSDCon 2026 Differential Revision: https://reviews.freebsd.org/D59610 MFC after: 3 days
This fixes the build with gcc 16's aggressive inlining. Reviewed by: ngie MFC after: 3 days Sponsored by: The FreeBSD Foundation Differential Revision: https://reviews.freebsd.org/D59538
This new manpage describes what the apple_bce driver services, how the driver is configured, and what Apple models the driver claims to support. Reviewed by: seuros Differential Revision: https://reviews.freebsd.org/D59467
sysrc: make independant from bsdconfig(8) MFC After: 1 week Approved by: dteske Reviewed by: dteske Differential Revision: https://reviews.freebsd.org/D59658
Remove source_rc_confs, inlined from defaults/rc.conf Inlining source_rc_confs broke sysrc_test:A_flag While here, switch to SPDX, bump version/copyrights, fixup comments. Prevent common function override when bsdconfig includes new sysrc.subr that contains inlined commons. Fix non-unique duplicate header variable incorrectly shared between bsdconfig's sysrc.subr include and new sysrc.subr include. Harden RC_DEFAULTS parameter expansion from DoS-via-glob (SC2223). Move pgm to the correct location to not override bsdconfig's pgm. Drop _SYSRC_JAILED=1 that is no-longer needed. Drop i18n from jail_depend now that messages are inlined to subr. Reported by: siva Fixes: https://cgit.freebsd.org/src/commit/?id=eaeb5f29bc6f sysrc: make independant from bsdconfig(8)
The passthrough ioctl has a userland ABI header, but nothing installed it. A program that wanted to use the ioctl had to copy the headers out of the source tree by hand. Install ufshci.h and ufshci_ioctl.h under /usr/include/dev/ufshci, the way nvme installs nvme.h. The ioctl header pulls in ufshci.h, so both go. Add the directory to the include mtree so installworld creates it. Reviewed by: imp (mentor) Sponsored by: Samsung Electronics Differential Revision: https://reviews.freebsd.org/D59560
For consistency, create a symbolic link from enic.4 to also if_enic.4 Reviewed by: ziaee, #manpages Differential Revision: https://reviews.freebsd.org/D59755 MFC after: 3 days
This warning triggers in a few places in contributed code, and would therefore be annoying to fix. Use -Wno-error= to at least show the warnings so there is some incentive to submit them upstream. MFC after: 3 days
Using a backtick inside a backtick-expansion isn't a good idea; switch instead to using $( ... ). Fixes: https://cgit.freebsd.org/src/commit/?id=a9710349513f ("EC2: Add desktop flavour") Sponsored by: Amazon
We don't have any packages to install by default when the instance first boots, so (as with the "small" flavour) don't enable the firstboot_pkgs script. If someone wants to launch a desktop with packages autoinstalled at first boot, they can enable the script at the same time as they provide the list of packages. Sponsored by: Amazon
When creating VM images, we filter the METALOG file created by pkg(8) when installing non-base packages, rejecting any lines which correspond to files which don't exist; this solves a problem which arose when a package was installed and then deinstalled (or upgraded) later in the image-building process. Unfortunately [ -e ... ] follows symlinks and is not basedir-aware, so an absolute symlink which is valid *inside* the image is omitted from the image if it points to something which isn't present in the build host system. Replace [ -e ... ] with [ -e ... ] || [ -L ... ] so that symlinks are included even if dangling. While I'm here, add quoting in case future paths become problematic. Sponsored by: Amazon MFC after: 3 days
Prior to this change, Dtrace and ZFS tests and their respective
directories would remain installed even if `${MK_DTRACE_TESTS}` == no or
`${MK_ZFS_TESTS}` == no.
This change enhances the logic to remove the tests when the respective
knobs are disabled so the tests will be removed when `make delete-old`
is run.
MFC after: 2 weeks
Prior to this change Dtrace's tests could be enabled even if
`${MK_DTRACE}` == no. Allowing this doesn't make sense so remove the
tests if `${MK_DTRACE}` is also disabled.
MFC after: 2 weeks
In the event someone specified MK_CDDL == no prior to this change, all
tests under `/usr/tests/cddl` would stick around after the change,
resulting in errors. Similarly, MK_DTRACE == no removed ctfconvert,
which the ctfconvert tests relied on (but remained around after the knob
was disabled).
Remove the tests under the respective blocks now so they don't linger
around on systems where these knobs are disabled, resulting in false
positives when running the complete test suite.
The ctfconvert tests weren't specifically put under
`${MK_DTRACE_TESTS}` == no because ctfconvert was bundled separately
from the upstream provided Dtrace test suite.
MFC after: 2 weeks
MFC with: 30c20ca26 ("Remove Dtrace/ZFS tests if their knobs are disabled")
This was added in the initial import of dtrace and overrides the normal load/unload targets with a custom target that loads a hardcoded set of modules. Over time, the set of modules has not been updated and is now incomplete. It's also not really useful compared to the default implementation of these targets used for loading or unloading an individual module being actively developed. This functionality is also available via dtraceall.ko which is how users commonly load the full suite of dtrace modules. Reviewed by: imp, markj Differential Revision: https://reviews.freebsd.org/D59821
The Mips architecture has been removed from all supported branches now. MFC after: 1 week
These aren't referenced in the tree anymore, so remove them from here. This doesn't chanage src.conf.5, so I didn't commit that file. Sponsored by: Netflix
Note there is ongoing work to add LoongArch support to the base system, but having target support in llvm is an essential component. This must be explicitly enabled using WITH_LLVM_TARGET_LOONGARCH. Reviewed by: dim MFC after: 1 week Differential Revision: https://reviews.freebsd.org/D59899
This completes step 5 from Committer's Guide. Approved by: jbo (mentor) Differential Revision: https://reviews.freebsd.org/D57934
In order to merge merge commits (such as vendor imports), we need to tell git cherry-pick which of the two branches referenced in the commit is the mainline. In our case, it is always the first. Approved by: markj
Previously we searched commits based on the author email address, but this isn't really right: if I commit something from a contributor, I'm still responsible for MFCing it, so really we should be filtering on the committer. Add a new --committer option to filter results by committer email address, defaulting to the user.email value in the git config. Keep the --author option, but don't filter by author unless the option is explicitly specified. Reported by: des Reviewed by: des Differential Revision: https://reviews.freebsd.org/D58126
This allows one to resume from a conflict with a plain `git cherry-pick --continue`, whereas before one would have to re-run the original git-mfc command after resolving the conflict and running `git cherry-pick --continue`. Suggested by: des Reviewed by: des Differential Revision: https://reviews.freebsd.org/D58129
Commit hashes listed in ~/.git-mfc-ignore are not listed in output of git-mfc --dangling or --pending. This is handy for silencing output about commits that are tagged for MFC or as fixing another commit, but which were not MFCed for some reason or other. Requested by: des Reviewed by: des Differential Revision: https://reviews.freebsd.org/D58161
This silences warnings when running git-mfc --pending against stable/13, 14 and 15. Reviewed by: des Differential Revision: https://reviews.freebsd.org/D58162
Sponsored by: Netflix
git-mfc: Slightly relax the regex used to search for reverts Prompted by commit 9dfaf1cb37f8ac89cf in FreeBSD src. Reported by: des
git-mfc: Let the upstream for PRERELEASE branches be main Such branches are in code slush but are the same as stable branches for the purpose of MFCs.
git-mfc: Improve handling of remotes If we can't figure out which remote to use, print a useful error instead of assuming that "freebsd" is the right remote to use.
- Make it work even when git arc isn't run from the root of the repo. - If the patch fails to apply, let git partially apply the patch and generate rej files for inspection. While here, remove the return value from apply_rev(), it's never actually used. Reviewed by: jhb Differential Revision: https://reviews.freebsd.org/D58532
Reviewed by: markj Differential Revision: https://reviews.freebsd.org/D58515
Somehow a few commits ended up with "null" appended to Nick's name and email address. Reviewed by: Nick Price <nick@spun.io> Differential Revision: https://reviews.freebsd.org/D58517
Instead of making the user run the underlying git-cherry-pick command after a conflict. Requested by: des Reviewed by: des Differential Revision: https://reviews.freebsd.org/D58514
Completed steps 5-6 and 10 in the committer's guide. Reviewed by: jhb Approved by: jhb (mentor) Differential Revision: https://reviews.freebsd.org/D58507
Phabricator user names are not useful identifiers outside of phabricator, don't use them if we can avoid it.
Reviewed by: dteske, fuz Approved by: dteske (mentor), fuz (mentor) Differential Revision: https://reviews.freebsd.org/D58700
Reviewed by: dteske, fuz Approved by: dteske (mentor), fuz (mentor) Differential Revision: https://reviews.freebsd.org/D58700
Reviewed by: dteske, fuz Approved by: dteske (mentor), fuz (mentor) Differential Revision: https://reviews.freebsd.org/D58700
There is no timeout period for merging from stable to releng branches, so we should ignore "MFC after" tags. While here, lift some uses of re.compile() out of loops. Reported by: des
In my repos the remote "freebsd" is git@gitrepo.freebsd.org:src.git. In particular, the last component is delimited by a colon, not a slash. Sponsored by: Klara, Inc. Differential Revision: https://reviews.freebsd.org/D48951
Show the differences between local commits and their associated Phabricator reviews, i.e., what "git arc update" would upload. For each commit, the review's current raw diff is applied to the commit's parent in a temporary index and the resulting tree is compared against the commit itself. An empty diff means the commit and the review are in sync. This makes it easy to check whether local amendments have diverged from the posted review before updating it, or to confirm that a review is current before landing. Reviewed by: markj Differential Revision: https://reviews.freebsd.org/D58789
And clean up luacheck warnings. Sponsored by: Klara, Inc. Differential Revision: https://reviews.freebsd.org/D48950
The header files for dialog, figpar, dpv were never listed. PR: https://bugs.freebsd.org/bugzilla/show_bug.cgi?id=297612 Fixes: https://cgit.freebsd.org/src/commit/?id=af202a5052b6 ("Retire dialog")
Add -t tag[,...] so a review can be tagged at creation instead of needing the web UI. Spaces in tag names are written as underscores; a leading # is optional. Reviewed by: markj Differential Revision: https://reviews.freebsd.org/D59019
git-sh-setup treats -h as help against an empty USAGE, so "git arc create -h" prints "usage: git arc". Handle -h before sourcing it so every subcommand prints the real synopsis. The create, stage, and update synopses showed optional commit-refs while git-arc(1) and the code require them. Advertise -p parent on create; the option was already implemented and documented. Sort create sub-command option-arguments alphabetically in three places: (1) synopsis from tool, (2) man-page synopsis, and (3) man-page description. Check for jq(1) / arc after checking for usage so -h always works. While here, fix missing "local o" in gitarc__stage(). Reviewed by: markj Differential Revision: https://reviews.freebsd.org/D59129
Reviewed by: markj Differential Revision: https://reviews.freebsd.org/D59162
Quote LOCALBASE and ARC_CMD default assignments so a poisoned value cannot glob into :'s argv. Reviewed by: markj Differential Revision: https://reviews.freebsd.org/D59130
I have been the de-facto maintainer for a few years now. Also mention the #scheduler group for reviews on Phabricator. Sponsored by: The FreeBSD Foundation
These files are not about scheduling per se but about priority management, mostly from user space.
Suggested by: des
Sponsored by: Chelsio Communications
Move back to the current committers section and record fuz and jrm as mentors. Approved by: portmgr Reviewed by: fuz (mentor), jrm (co-mentor) Differential Revision: https://reviews.freebsd.org/D59500
Approved by: adrian (mentor) Reviewed by: adrian Differential Revision: https://reviews.freebsd.org/D59860 Signed-off-by: Nick Price <nprice@FreeBSD.org>
Reviewed by: adrian Differential Revision: https://reviews.freebsd.org/D59861
Add myself formally as the "maintainer" for atf, kyua, and lutok, both in the GitHub CODEOWNERS and MAINTAINERS files. This change matches the herald rules I recently setup for these third-party components.
Add myself as the maintainer for routing subsystem in the GitHub CODEOWNERS and MAINTAINERS files.
This was installed in two different places.
Sponsored by: The FreeBSD Foundation Differential Revision: https://reviews.freebsd.org/D59912
Approved by: imp (mentor), des (mentor) Differential Revision: https://reviews.freebsd.org/D59871
Approved by: imp (mentor), des (mentor) Differential Revision: https://reviews.freebsd.org/D59871
Approved by: imp (mentor), des (mentor) Differential Revision: https://reviews.freebsd.org/D59871
There is no functional change for existing tests, but allows to write a test that would expect an immediate success of bind(2).
- Test SOCK_DGRAM (UDP) sockets. - Test binding to 0:port and to a addr:port in presence of connected socket using the port. Differential Revision: https://reviews.freebsd.org/D56707
Just avoid repeating the test program name in every test case name. No functional change. Reviewed by: markj Differential Revision: https://reviews.freebsd.org/D56727
This fixes an endianness bug in sys/netinet/ip_reass_test. Just use the code from RFC 1071. Reported by: glebius Reviewed by: glebius, Timo Völker MFC after: 1 week Sponsored by: Netflix, Inc. Differential Revision: https://reviews.freebsd.org/D57988
/sbin/ping and /sbin/ping6 are hard-linked, and the vmmap sysctl handler doesn't know which name was used to launch the process. PR: https://bugs.freebsd.org/bugzilla/show_bug.cgi?id=296116 MFC after: 3 days Fixes: https://cgit.freebsd.org/src/commit/?id=080a4087014e ("tests: Fix race condition in aslr_setuid")
MFC after: 1 week
Reviewed by: kib
MFC with: 5c32aa785184 ("kern: add pdopenpid(2)")
Differential Revision: https://reviews.freebsd.org/D58023
Reviewed by: markj Sponsored by: The FreeBSD Foundation MFC after: 1 week Differential revision: https://reviews.freebsd.org/D57163
In order to reuse the sendfile_helper program in pf tests, move it to tests/sys/common directory, indicatint that it is also used from another places than sys/kern. Also make the readlen variable static. Reviewed by: gelbius, kp Differential Revision: https://reviews.freebsd.org/D58040
In order to use the sendfile_helper program in a pf test script that requires non-loopback interfaces, add functionality to sendfile with a TCP socket that is connected to a remote host. The behavior for unix sockets and TCP loopback sockets is unchanged. Reviewed by: glebius Differential Revision: https://reviews.freebsd.org/D58041
MFC after: 3 days
Reported by: gcc -Werror=shadow Reviewed by: asomers, markj Fixes: https://cgit.freebsd.org/src/commit/?id=ee1c3d38a26a ("fusefs: fix vnode locking violations during execve") Differential Revision: https://reviews.freebsd.org/D58130
Make sure we have reachability when one of our nexthops gets down without deleting the route. Differential Revision: https://reviews.freebsd.org/D57552
This is a script that eventually will test boot with qemu all the supproted combinations for the boot loader. There's several things that could be done with gptboot or boot0sio (or not) that aren't tested. We don't test the 10-odd hardware root devices we support, nor do we test complex scenarios like RELAXED vs STRICT zfs efi booting. However, the scenarios we do support are included here. We test aarch64, amd64, armv7, powerpc64, powerpc64le, and riscv64 for BIOS, UEFI, and Prep and OpenFirmware (as appropriate) crossed with CDROM, MBR and GPT (and some hybrid) crossed with lua, 4th and simple loaders. Plus some linuxboot and memdisk scenarios, including the recently added compression for ram disk scenarios: === Results: 67 passed, 3 failed, 9 timed out (of 79) === The timeouts are well understood, usually failure to find the root disk. The failures are bad console assumptions. netboot-bios fails because TFTP with a single packet buffer in qemu gives horrible throughput, so the test takes 18-20 minutes. Now that I have a dashboard, I can fix the rest one by one. There's also a powerpc architecture that you can request specifically, but it's just for convenience and tests with the non-functional mac99 qemu machine. I will eventually eliminate this architecture. I added it to make sure the FreeBSD version wasn't too hard coded since this framework pulls from CD images to get the binaries for the minimal root used in testing and there's no 15.x 32-bit powerpc images. We need to add http and nfs root booting tests, but that's for the future. Plus there's some other functional tests that we should also add for different types of root (usb, sata, sas, nvme, ufs, emmc, sd, etc) that would be useful to test, especailly the non-sata/non-nvme ones. How we do that is still TBD. I leaned on claude to iterate over the recipes that I've developed over the years, collected off the internet or got on IRC recently to produce this framework. Most of this code is fairly good, while a few parts, especailly some of the comments, are detectable as AI produced. My plans are to iteratively improve those. Since this is just a test, and since I've broken many scenarios w/o realizing, it's a good tradeoff. I've not made it an ATF test since we test all the architectures, but I'm open to feedback in this area. Total time to test all the architectures is about 10 minutes. It assumes you've built GENERIC* and the boot loader for all the architectures too. In the future, I plan on moving to MINIMAL for all the boot testing, but likely only after PCI devmatch is integrated into it. That would be incrementally faster test times. The man page is decent, but was also generated by Claude with only trivial edits by me to date.... But at least there's a man page for it, though neither it nor the script is installed onto the system. Sponsored by: Netflix Assisted-by: Claude Code (Opus 4.6, Opus 4.8(1M) and Sonet 5.0) Differential Revision: https://reviews.freebsd.org/D58008
Help validate my assertion that "physmem will never report empty ranges". Part of this is covered by the existing tests, which check the merging of adjacent/overlapping regions. The other part is to ensure that addition of zero-sized ranges is ignored. The physmem implementation also includes logic to ignore the first physical page of memory (physical addresses 0 to PAGE_SIZE-1). Add a second test case for this. Reviewed by: markj MFC after: 1 week Differential Revision: https://reviews.freebsd.org/D45914
The fix for this is being tracked upstream here: https://github.com/onetrueawk/awk/issues/269 While here, just cd into $SRCDIR while executing tests, since the test engine isolates every testcase's working directory. This ensures that the xfail actually applies to the next command. Reviewed by: mhorne MFC after: 3 days Sponsored by: The FreeBSD Foundation
Attach a process-mode counting PMC to the current process, start it, then detach and release it while it is still loaded on the hardware - the case that previously leaked the PMC's runcount reference and wedged pmc_wait_for_pmc_idle() at release. A second case does the same from a multi-threaded process so the sibling threads' references have to be drained too. The tests need an allocatable process-mode counting event and skip where none is available (hwpmc(4) not loaded, or a VM without a vPMU). Reviewed by: adrian MFC after: 2 weeks Assisted-by: Claude Code (Fable 5) Differential Revision: https://reviews.freebsd.org/D58343
stress2: Updated the exclude file
stress2: Added a comment
stress2: Added a regression test
Skip the message-content check on kernels that do not advertise the exterr_strings feature, and pin the output format by clearing EXTERROR_VERBOSE. Reviewed by: kib MFC after: 1 week Assisted-by: Claude Code (Fable 5) Differential Revision: https://reviews.freebsd.org/D58322
This requested fix[0] was not complete before the change was committed. Cleans up this error message when running tests[1]: "Cannot 'start' ipfilter. Set ipfilter_enable to YES in /etc/rc.conf or use 'onestart' instead of 'start'." [0] https://reviews.freebsd.org/D21065?id=60288#inline-131488 [1] https://ci.freebsd.org/job/FreeBSD-main-amd64-test/28917/testReport/sys.netpfil.common/rdr/ipfnat_local_redirect/ Fixes: https://cgit.freebsd.org/src/commit/?id=f97a8a36153a9 MFC after: 3 days Sponsored by: The FreeBSD Foundation
This keeps the skipped test message consistent with others. Reviewed by: netchild MFC after: 3 days Sponsored by: The FreeBSD Foundation
The time_unit test case uses PID 1 as a target for pwait. This doesn't work in a jail. Since all we need is a process that we know won't die while the test is running, we may as well use ourselves. MFC after: 1 week Sponsored by: Klara, Inc. Sponsored by: NetApp, Inc. Reviewed by: ngie Differential Revision: https://reviews.freebsd.org/D58418
The function create_staticobj() is only used inside this translation unit. Clang produces a -Wmissing-prototypes warning during standard buildworld. This warning will become a fatal compile error if MK_WERROR is enabled for hardened builds. PR: https://bugs.freebsd.org/bugzilla/show_bug.cgi?id=285870 Fixes: https://cgit.freebsd.org/src/commit/?id=ee9ce1078 ("libc: tests: add some tests for __cxa_atexit...") Signed-off-by: Zhang Qiyue <peter-open-source.probing805@aleeas.com> Reviewed-by: ngie Pull-Request: https://github.com/freebsd/freebsd-src/pull/2321
This change converts the longhand form of `extern "C" {` and its
corresponding `}` into `__BEGIN_DECLS` and `__END_DECLS`, respectively.
The new form is much easier to grep for and is a best practice to use in
the FreeBSD tree.
This is meant to be a non-functional change.
MFC after: 1 week
pwait: Test the new -r option Test that pwait without -r reports a process as soon as it terminates, while pwait with -r does not report it until it has been reaped. MFC after: 1 week Sponsored by: Klara, Inc. Sponsored by: NetApp, Inc. Reviewed by: kib Differential Revision: https://reviews.freebsd.org/D58385
pwait: Fix pwait_normal test case Reported by: markj Fixes: https://cgit.freebsd.org/src/commit/?id=e115066370dc ("pwait: Test the new -r option")
Add tests for both IPv4 and IPv6 routes with the prefsrc attribute. Also test IPv4 routes over IPv6 nexthops and borrow their IPv4 addresses from the loopback interface. Reviewed by: glebius Differential Revision: https://reviews.freebsd.org/D58326
The child exited immediately after pdfork(), so the parent's pdopenpid() could catch it mid-exit (P_WEXIT) and fail with EBUSY. Block the child on a pipe until the parent has opened the second descriptor, then release it Approved by: markj Sponsored by: Netflix Differential Revision: https://reviews.freebsd.org/D58546
Add a regression test for gre(4) to make sure all of the gre capabilities and options are working as intended. Reviewed by: markj Differential Revision: https://reviews.freebsd.org/D55363
Some of the preadv() and readv() tests were not initializing the iovecs they pass to the system call. When the system call is expected to fail, that's fine since the FORTIFY_SOURCE checks cause the process to be aborted. However, in the rest of the test cases, the (p)readv() call could cause spurious test failures, e.g., when an uninitialized iov entry points to the current stack frame and the canary gets overwritten. Modify the tests to explicitly initialize iov entries to avoid this. The "iov" variants don't have this problem, so leave them alone. Reviewed by: kevans MFC after: 2 weeks Sponsored by: The FreeBSD Foundation Differential Revision: https://reviews.freebsd.org/D58289
Sponsored by: The FreeBSD Foundation MFC after: 1 week Differential revision: https://reviews.freebsd.org/D58530
Reviewed by: markj Sponsored by: The FreeBSD Foundation MFC after: 1 week Differential revision: https://reviews.freebsd.org/D58569
pdwait's capsicum/enotcap and procdesc's pdopenpid_capmode enter capability mode. Require security_capability_mode (and security_capabilities for enotcap) so the cases skip cleanly on kernels built without CAPABILITIES instead of failing. Approved by: asomers, gallatin Sponsored by: Netflix Differential Revision: https://reviews.freebsd.org/D58545
Sponsored by: The FreeBSD Foundation MFC after: 1 week
Cover the new fd-direct connect path: stream connect and data passing, the peer address reported by `getpeername(2)`, datagram to an unbound peer, the `EINVAL`/`ENOTSOCK`/`EPROTOTYPE`/`ECONNREFUSED` error matrix, and the Capsicum token semantics — a descriptor limited to `CAP_CONNECTAT` can be connected to but not listened on, accepted from, or read, and one lacking `CAP_CONNECTAT` cannot be a connect target. Stream listeners are always bound: `uipc_listen()` refuses unbound sockets with `EDESTADDRREQ`, so an unbound fd-direct listener is not reachable even with this feature. Signed-off-by: John Ericson <John.Ericson@Obsidian.Systems> Assisted-by: Claude Code (Claude Opus 4.8 and Fable 5) Reviewed by: markj MFC after: 2 months Differential Revision: https://reviews.freebsd.org/D58406
This test was already marked as always skipped. Fixes: https://cgit.freebsd.org/src/commit/?id=069a67374ed9641ff1ada2aecaac1cc61a560649 Reviewed by: pouria Differential Revision: https://reviews.freebsd.org/D58114
tests/netinet/socket_afinet: unroll multibind test into a table The test has 6 dimensions: address family, socket type, socket option on the first socket, socket option on the second socket, is first socket bound to specific address or wildcard and is the second socket priveleged or not. Before the change 3 dimensions are implemented as 3 nested for() loops, 2 dimensions are implemented as repetitions in the test body and one dimension as two actions in the innermost loop. I'm about to add one more dimension: whether the second socket is bound to a specific address or wildcard instead of using first socket's getsockopt(2) result. Also, there is a change under discussion that would make SOCK_STREAM sockets behave different to SOCK_DGRAM. That would break result consistency in the dimensions of socket type. We expect that consistency in the dimension of address families shall never break, thus this one remains a for() loop. The priveleged & non- privileged bind(2) attempts also remain as two actions, but expected results are in the table. The rest of dimensions are unrolled into a table, which at the moment has quite a lot of lines with identical results. However, as more tests are added and SOCK_STREAM behavior changes, the table will get more mixed results. Also, reading a test that is written in a declarative manner a table is much easier and modifying it is more resistent to accidential breakage. Differential Revision: https://reviews.freebsd.org/D58085
tests/netinet/socket_afinet: multibind second socket can be different Allows to add tests to the table where the second socket doesn't take address from the first. No functional change yet, all tests test the same conditions. Differential Revision: https://reviews.freebsd.org/D58087
tests/netinet/socket_afinet: add more tests to multibind Add tests where first socket and second socket are bound to different addresses, e.g. first specific and second wildcard and vice versa. Mark success with SO_REUSEPORT on the second socket as a bug suspect. Mark failure to bind to INADDR_ANY in presence of other UID's specific bound socket to the same port as probably too strict. Differential Revision: https://reviews.freebsd.org/D58088
Fix a typo in the ZFS multi_dataset_4 test, where a path separator was missing. Reported by: markj MFC after: 1 week
Reviewed by: fuz Approved by: fuz (mentor) MFC after: 1 month Pull Request: https://github.com/freebsd/freebsd-src/pull/2352
We created so many states that our bulk-sync occasionally caused epair to drop packets, which in turn caused the test to fail. That's not what we're testing here, make it more robust by creating fewer states. PR: https://bugs.freebsd.org/bugzilla/show_bug.cgi?id=297307 Sponsored by: Rubicon Communications, LLC ("Netgate")
Add three test cases verifying fts(3) Capsicum capability mode: - fts_dirfd_valid: verifies fts_dirfd is set for all non-root entries and openat(fts_dirfd, fts_name) identifies the same inode as fts_accpath - fts_dirfd_capsicum: verifies complete fts traversal works in Capsicum capability mode using fts_openat() and fts_dirfd - fts_dirfd_deep_tree: verifies fts_dirfd + fts_name is correct at all directory depths (7 non-root entries) Sponsored by: Google LLC (GSoC 2026) Reviewed by: asomers Pull Request: https://github.com/freebsd/freebsd-src/pull/2332
This is enough to show that the idea works (and to exercise uexterr_set()), but isn't complete by any means. Reviewed by: kib Effort: CHERI upstreaming Sponsored by: Innovate UK Differential Revision: https://reviews.freebsd.org/D58060
Verify that unsupported, unterminated, and overly long formats output expected messages. Reviewed by: kib Effort: CHERI upstreaming Sponsored by: Innovate UK Differential Revision: https://reviews.freebsd.org/D58413
On FAT file systems, access time has a resolution of 1 day, so it is really the access date. Strip the time component from the epoch timestamp in order to check the access time. Reference: https://learn.microsoft.com/en-us/windows/win32/sysinfo/file-times Reviewed by: ngie MFC after: 1 week Differential Revision: https://reviews.freebsd.org/D54584
The bash-ism here results in an "ambiguous output redirect" error when run under tcsh. Similar unionfs stress2 tests don't do this, and unionfs7 doesn't seem to generate spurious output when run locally, so just delete it. Reviewed by: pho Differential Revision: https://reviews.freebsd.org/D58856
This introduces the requirement preparation handler concept, with the
first handler implemented to load the required kernel modules.
Running without arguments lists all available handlers:
kyua prepare
Currently there are only two handlers:
all: runs all available handlers
kmods: loads the modules declared by required_kmods metadata
The dry run option lists the required modules without actual loading:
kyua prepare { --dry-run | -n } kmods
The "prepare" command traverses only the given tree of tests, i.e., the
following invocation lists all required modules for the whole test
suite:
kyua prepare -k /usr/tests/Kyuafile -n kmods
, while this one is limited to the pf tests only:
kyua prepare -k /usr/tests/sys/netpfil/pf/Kyuafile -n kmods
Reviewed by: ngie
Differential Revision: https://reviews.freebsd.org/D48087
Suggested by: asomers
This generalizes the divapp logic and makes it usable for other test scenarios such as diverted TCP connections. Reviewed by: markj Differential Revision: https://reviews.freebsd.org/D59067
This tests what FreeBSD-SA-26:56.hwpmc fixed. exec_setgid_drops_pmc asserts the kernel takes a process-mode PMC away when its target execs a set-gid program its owner is not entitled to trace. exec_setuid_no_double_unlink lets the target exec a set-uid program; the teardown must unlink the process descriptor exactly once, and completing at all is the assertion. Both need an unprivileged owner and must not drop privileges themselves, since p_candebug() would then refuse the target to its own owner; they ask for require.user instead. MFC after: 1 month MFC to: stable/15 MFC to: stable/14 Assisted-by: Claude Code (Opus 5)
A pmc_id_t is a packed integer that the driver hands to userland and accepts back on eleven operations, and nothing tested what happens when one comes back forged, stale, or belonging to another process. Neither was there a test that an unprivileged caller is refused the operations that need a privilege. The cases use a SOFT-class PMC wherever the counter itself does not matter, so they run on a machine with no PMU. MFC after: 1 month MFC to: stable/15 MFC to: stable/14 Assisted-by: Claude Code (Opus 5)
Seven ATF cases covering process-attachment teardown orderings: a target that exits before it is detached, the owner that exits before its target (hwpmc's other unlink path), releasing a still-running attached PMC, row exhaustion with out-of-order release, and PMC_F_DESCENDANTS inheritance including a fork storm. All pass on a debug (INVARIANTS+WITNESS) and a KASAN kernel. MFC after: 1 month MFC to: stable/15 MFC to: stable/14 Assisted-by: Claude Code (Opus 5)
The companion to pmc_exec_test.c, which covers only the drop side of a credential-changing exec. Three cases cover what the drop must not overreach into: an exec that changes no credentials keeps the PMC, a set-id exec whose credential change the kernel suppresses for a traced target keeps it too, and a set-id fexecve(2) drops it. They exercise the permission logic FreeBSD-SA-26:56.hwpmc reworked, not the defect it fixed. All three pass on a debug (INVARIANTS+WITNESS) kernel. The two keep-cases were each observed to fail on a kernel mutated to detach unconditionally. MFC after: 1 month MFC to: stable/15 MFC to: stable/14 Assisted-by: Claude Code (Opus 4.8)
Nine ATF cases covering PMC_OP_CONFIGURELOG and the descriptor-less log operations: which descriptors are accepted, when a log is required in the first place, and what the log operations do without one. MFC after: 1 month MFC to: stable/15 MFC to: stable/14 Assisted-by: Claude Code (Opus 5)
Reviewed by: markj Sponsored by: The FreeBSD Foundation MFC after: 1 week Differential revision: https://reviews.freebsd.org/D58586
Adjust test to check for ECAPMODE using grandchild instead of child. Childrens can be opened even in cap mode. Add test for the later. Reviewed by: markj Sponsored by: The FreeBSD Foundation MFC after: 1 week Differential revision: https://reviews.freebsd.org/D58989
Do not assign `b` in the `ATF_REQUIRE` macro. Set and test `b` separately to avoid the issue cited by clang++/g++ after implementing the change referenced in [1]. MFC after: 2 weeks Reported by: clang (-Wparenthesis) Reference: https://github.com/freebsd/atf/pull/72
kqueue returns a value != -1 on error. Test for that instead of any non-zero value to confirm that success was achieved when calling `kqueue`. This issue exists with ATF 0.22+ [1]. MFC after: 2 weeks Reported by: clang (-Wparenthesis) [1]: https://github.com/freebsd/atf/pull/72
The variable is written in the child process which shares the address space with the parent. The data flow must not be optimized by a compiler. Reviewed by: markj Fixes: https://cgit.freebsd.org/src/commit/?id=ddf62c83fc0a ("sys/tests/kern/pdopenpid: pdopenpid(2) is allowed in cap mode") Sponsored by: The FreeBSD Foundation MFC after: 1 week Differential revision: https://reviews.freebsd.org/D59282
Cast the size_t quantity used in a comparison to off_t to mute a `-Wsign-compare` complaint that now occurs after ATF 0.22 [1]. MFC after: 2 weeks Reported by: clang Reviewed by: markj Differential Revision: https://reviews.freebsd.org/D59285 [1]: https://github.com/freebsd/atf/pull/72
Confirm that creating clients/sockets was successful by testing the result separate from the assignment and test that the return value is not -1 instead of testing that the value returned is non-zero. This fixes the build with [ATF 0.22+][1]. MFC after: 2 weeks Reported by: clang (-Wparenthesis) Reviewed by: tuexen, cc Differential Revision: https://reviews.freebsd.org/D59284 [1]: https://github.com/freebsd/atf/pull/72
PR: https://bugs.freebsd.org/bugzilla/show_bug.cgi?id=296944 Reviewed by: cy MFC after: 3 days MFC to: stable/15 Sponsored by: The FreeBSD Foundation
Previously, we would precompile D test dependencies using the host's dtrace, which unconditionally outputs ELF files in the host's format. This breaks the cross-compile build with errors like the following: dtrace: failed to link script: incorrect ELF machine type for object file: tst.usdt.pieo --- usdt.o --- *** Failed target: usdt.o This patch moves compilation to runtime for all C-based testcases that have a dependent D source file. Reviewed by: markj MFC after: 1 week Sponsored by: The FreeBSD Foundation Differential Revision: https://reviews.freebsd.org/D59030
This serves to catch the regression fixed by commit
1a669b66ddb4 ("syslogd: reap pipe children on config reload").
MFC after: 1 week
Exercise each of the 32-bit BIOS Parameter Block and FSInfo fields that readboot() decodes, using values whose most significant byte has its high bit set. Each case checks two things: that fsck_msdosfs(8) reports the full unsigned 32-bit value back on stdout, and that nothing writes a sanitizer runtime error to stderr. The second check is what catches a byte-at-a-time decode. Shifting such a byte left by 24 is undefined, but every compiler we use wraps it into the same bit pattern, so the decoded value alone cannot tell a correct decode from an overflowing one. In a WITH_UBSAN build bsd.sanitizer.mk compiles with -fsanitize=undefined and -fsanitize-recover=undefined, so the shift is reported on stderr and execution continues, which the test can then assert on. Against the byte-at-a-time decode these cases fail in a WITH_UBSAN build and pass without it. Note that the stderr check also fails on unrelated undefined behavior that these images reach anywhere in fsck_msdosfs(8), which is intended. MFC after: 1 week
Build a 4.5 GiB FAT32 image whose LOST.DIR cluster sits exactly 4 GiB above the single cluster of a PAYLOAD.BIN, so that truncating the offset of the former to 32 bits yields the offset of the latter, then inject a lost cluster chain and let fsck_msdosfs(8) reconnect it. The test asserts both halves of the bug fixed in the previous commit: that PAYLOAD.BIN's cluster is unchanged, and that a second pass no longer reports the chain as lost, which it only stops doing once the directory entry reaches the real LOST.DIR. newfs_msdos(8) -C only calls ftruncate(2) and nothing outside the reserved area, the FATs and a handful of clusters is ever written, so the image stays sparse and costs about 2 MiB on disk. The geometry is read back out of the BPB rather than assumed, so newfs_msdos(8) stays free to lay the file system out differently; the test fails with a clear message if the volume ever becomes too small to hold a cluster a full 4 GiB beyond the data area. MFC after: 1 week
hwpmc: Add ATF regression tests for hwpmc EXTERROR diagnostics
Root-only ATF program hitting negative allocate/attach/read-write paths
and asserting the exterr(3) text. AMD/IBS cases skip without the PMC
class; program skips without hwpmc.
Additional changes by mhorne@:
- Move and rename to the established test directory tests/sys/pmc
- Remove broken test amd_missing_pmu_flag; fixed by recent change
6c4d9b9af1a3
- Add ATF_REQUIRE_FEATURE("exterr_strings") to skip the tests on kernels
compiled without the strings
- Remove arch-conditional compilation; tests are properly gated by PMC
class check
- Fix copyright formatting
Reviewed by: Ali Mashtizadeh <ali@mashtizadeh.com>
Signed-off-by: Andre Silva <andasilv@amd.com>
Co-authored-by: mhorne
Sponsored by: AMD
Pull Request: https://github.com/freebsd/freebsd-src/pull/2180
pmc_exterr_test: add MACHINE_ARCH check The tests manipulate MD structure fields directly. Build these tests for amd64 only. Fixes: https://cgit.freebsd.org/src/commit/?id=e555692d1bb ("hwpmc: Add ATF regression tests for hwpmc EXTERROR diagnostics")
The previous pattern is cited as an issue with ATF 0.22+ when using clang/gcc after [1]. MFC after: 2 weeks Reported by: clang (-Wparenthesis) Differential Revision: https://reviews.freebsd.org/D59274 [1]: https://github.com/freebsd/atf/pull/72
Code that assigned variables as part of ATF_\* are no longer permitted due to changes introduced in [ATF 0.22][1]. MFC after: 2 weeks Reported by: clang/gcc (-Wparenthesis) Differential Revision: https://reviews.freebsd.org/D59286 [1]: https://github.com/freebsd/atf/pull/72
Check that wildcard bind(2) doesn't yet select a local address, and later connect(2) succeeds.
Reviewed by: asomers MFC after: 3 days Sponsored by: The FreeBSD Foundation
This fixes the build with gcc 16 after the changes to -Wunused*[0]. [0] https://gcc.gnu.org/gcc-16/porting_to.html#changes-to-wunused Reviewed by: markj MFC after: 3 days Sponsored by: The FreeBSD Foundation Differential Revision: https://reviews.freebsd.org/D59542
These exercise lookup traversal from tmpfs into unionfs, which is useful because, unlike UFS, tmpfs does not allow recursion on its vnode locks by default. unionfs22.sh exercises these lookups with a normal unionfs mount, while unionfs23.sh uses '-o below' for the unionfs mount and reproduces the panic described in PR 298201. Reviewed by: kib, markj, pho Tested by: pho Differential Revision: https://reviews.freebsd.org/D59494
The script was passing the wrong device name to `graid3 insert` and didn't notice that the command was failing. MFC after: 1 week Event: EuroBSDcon 2026 DevSummit Reviewed by: delphij Differential Revision: https://reviews.freebsd.org/D59564
Reviewed by: dteske, fuz Approved by: dteske (mentor), fuz (mentor) MFC after: 1 week Differential Revision: https://reviews.freebsd.org/D59611
Split accepted ABIs test into separate case and conditionally xfail it on archs that don't have linux(4) support. This fixes a CI test failure[0]. [0] https://ci.freebsd.org/job/FreeBSD-main-riscv64-test/16706/testReport/usr.bin.truss/truss_test/unknown_syscall/ Reviewed by: dteske MFC after: 3 days Sponsored by: The FreeBSD Foundation Differential Revision: https://reviews.freebsd.org/D59594
Helps to mitigate failures due to timeout in CI on emulated architectures: https://ci.freebsd.org/job/FreeBSD-main-riscv64-test/16565/testReport/usr.bin.gh-bc/ MFC after: 3 days Sponsored by: The FreeBSD Foundation Differential Revision: https://reviews.freebsd.org/D59547
This avoids globally set timeouts for the group of tarfs_large tests on slower emulated architectures. While here, lower each testcase's timeout to reflect the reduction of work. On QEMU aarch64, the largest case runs in ~500s on a modern desktop, so double that for a conservative estimate. Discussed with: des MFC after: 3 days Sponsored by: The FreeBSD Foundation Differential Revision: https://reviews.freebsd.org/D59033
Test that ipfw's layer-2 hook copes with unmapped mbufs, on the pass and on the deny path. Reviewed by: glebius Assisted-by: Claude Code (Fable 5, Opus 5) Differential Revision: https://reviews.freebsd.org/D59390
unix_connectat's fdescfs cases call mount_fdescfs(), which skipped on ENODEV. nmount(2) never returns ENODEV: vfs_donmount() remaps the ENODEV from a failed fdescfs module load to EINVAL with errmsg "Invalid fstype", so the skip never fired and the cases failed on kernels without fdescfs. Approved by: ngie, asomers Sponsored by: Netflix Differential Revision: https://reviews.freebsd.org/D59227
Load, list and clear the fingerprints.
Sponsored by: Rubicon Communications, LLC ("Netgate")
Pin the overlay namei(9) uses for Linux ABI processes: a target under the ABI root wins (PR 289739), a native-only target should retry from the native root (PR 297426), an ENOENT past a resolved ABI target is not retried, and the plain no-symlink native fallback is unchanged. Skip without a Linux userland so CI stays green. The two native-only cases use atf_expect_fail until the retry lands. PR: https://bugs.freebsd.org/bugzilla/show_bug.cgi?id=297426 Reviewed by: kib Differential Revision: https://reviews.freebsd.org/D59990
Cirrus CI shutted down its service on June 1, 2026 and its CI dashboard website (https://cirrus-ci.com) is now inaccessible. Since these files are no longer used, remove them from the repository. Reviewed by: brooks, emaste, lwhsu Approved by: olce (mentor) MFC after: 2 weeks Sponsored by: FreeBSD Foundation Differential Revision: https://reviews.freebsd.org/D59858
Approved by: nprice Sponsored by: Netflix Differential Revision: https://reviews.freebsd.org/D60145
Approved by: asomers Sponsored by: Netflix Differential Revision: https://reviews.freebsd.org/D60144
geom_subr.sh serves both legacy TAP tests and ATF tests. It defaults to the TAP path, so ATF tests must set ATF_TEST=true before sourcing it. Sponsored by: Netflix
This fixes the build with gcc 16 after the changes to -Wunused*[0]. [0] https://gcc.gnu.org/gcc-16/porting_to.html#changes-to-wunused Reviewed by: netchild MFC after: 3 days Sponsored by: The FreeBSD Foundation
These could go in other categories, but it's more clear if they're here instead.
The flag is -D, but it was written as a second -d. Add a period too. MFC after: 3 days
Reviewed by: markj Differential Revision: https://reviews.freebsd.org/D57928
Fixes: https://cgit.freebsd.org/src/commit/?id=c3c8f4d9e662 ("cpu: New cpu_get_pcpuid(), retrieves internal CPU ID") Sponsored by: The FreeBSD Foundation
Rename handler function type 'lapic_thermal_handle_function' to the
shorter 'lapic_thermal_handler_t'. Move it closer to the function
declaration block where it is used. Make it a true function type (no
pointer) and add explicit pointer marks on usage.
Rename 'lapic_thermal_function_value' to the more immediately clear
'lapic_thermal_function_arg'. In lapic_thermal_enable(), use 'func_arg'
as the argument name for the handler argument, which at least refers to
function 'func', rather than the generic 'value'.
Finally, rename the global handler variable from
'lapic_thermal_function_ptr' to the shorter 'lapic_thermal_function'
(dynamic functions can be referenced only through a pointer).
MFC with: 87ba088fa310 ("x86/local_apic.c: Add support for installing a thermal interrupt handler")
Sponsored by: The FreeBSD Foundation
MFC after: 1 week
We should exit the net epoch instead of acquiring one more entry. Reported by: Wafa Hamzah <wafah@nvidia.com> Reviewed by: jhb Sponsored by: Nvidia networking Fixes: https://cgit.freebsd.org/src/commit/?id=4726b80d9379 ("OFED: Various changes from Linux 4.20") Differential revision: https://reviews.freebsd.org/D58127
Reviewed by: mckusick Discussed with: markj Tested by: pho Sponsored by: The FreeBSD Foundation MFC after: 1 week Differential revision: https://reviews.freebsd.org/D57658
Sponsored by: The FreeBSD Foundation MFC after: 3 days
if administrator mistakenly types into configuration file anchor=authpf_test where 'authpf_test' is followed by white space, the authpf(8) is going to use anchor 'authpf_test ' instead of the 'authpf_test' which is defined in pf.conf(5) as 'anchor authpf_test/*' issue kindly reported and patch submitted by Avinash Duduskar <avinash.duduskar (_at_) gmail (_dot_) com> OK sashan@ PR: https://bugs.freebsd.org/bugzilla/show_bug.cgi?id=296958 MFC after: 1 week Obtained from: OpenBSD, sashan <sashan@openbsd.org>, 2d12a8e44d Sponsored by: Rubicon Communications, LLC ("Netgate")
Reviewed by: markj Sponsored by: The FreeBSD Foundation MFC after: 1 week Differential revision: https://reviews.freebsd.org/D58463
Also use bool for the 'cancel' argument for cond_wait_common(). Reviewed by: markj Sponsored by: The FreeBSD Foundation MFC after: 1 week Differential revision: https://reviews.freebsd.org/D58463
Show nhop flags like invalid nexthop to debug cases like the PR below. PR: https://bugs.freebsd.org/bugzilla/show_bug.cgi?id=296883 Reviewed by: glebius Differential Revision: https://reviews.freebsd.org/D58347
manuals: Fix Fx and nearby mechanical typos Fix compiler warnings related to the Fx macro, as well as all other mechanical typos that were visible within one screenful of them. These cause rendering glitches on various toolchains with various of the five and a half decades of rich output formats and tooling manpages scale to. The *x macro set specifies operating systems. These macros take the rest of the line as an argument. Sometimes, a space was not used to separate the argument of Fx and the trailing period. Another, FreeBSD Foundation was misrepresented as an operating system version instead of an author. Two more had other parts of the sentence supplied as an argument to Fx. While I had those open, fix the other mechancial typos visible on those specific screenfulls. Fix a list width glitch, correct section typo AUTHOR to AUTHORS, and switch AUTHORS sections containing prose to prose-mode so that they wrap freely when rendered. PR: https://bugs.freebsd.org/bugzilla/show_bug.cgi?id=297248 MFC after: 3 days Fixes: https://cgit.freebsd.org/src/commit/?id=d39e310c7d6a ("man/man3: add stdbit.3") Fixes: https://cgit.freebsd.org/src/commit/?id=d790b16bbf0c ("add man pages for stdbit functions") Fixes: https://cgit.freebsd.org/src/commit/?id=b61850c4e6f6 ("net.link.bridge.member_ifaddrs to false") Reported by: wosch (groff is complaining about incorrect Fx usage)
manuals: Fix more Fx and nearby mechanical typos Fix compiler warnings related to the Fx macro, as well as all other mechanical typos that were visible within one screenful of them. These cause rendering glitches on various toolchains with various of the five and a half decades of rich output formats and tooling manpages scale to. The *x macro set specifies operating systems. These macros take the rest of the line as an argument. Sometimes, a space was not used to separate the argument of Fx and the trailing period. Others had other parts of the sentence supplied as an argument to Fx. While here, fix the other mechanical typos visible on those specific screenfulls. Correct section typo AUTHOR to AUTHORS, markup utilities with Sy, and apply line break after the end of a sentence. PR: https://bugs.freebsd.org/bugzilla/show_bug.cgi?id=297248 MFC after: 3 days Reported by: wosch (are you sure that's all of the broken Fx'es?) Fixes: https://cgit.freebsd.org/src/commit/?id=ff2bc641599a ("Fix Fx and nearby mechanical typos") Fixes: https://cgit.freebsd.org/src/commit/?id=d790b16bbf0c ("add man pages for stdbit functions") Fixes: https://cgit.freebsd.org/src/commit/?id=6c57e368eb17 ("implement C23 memalignment()") Fixes: https://cgit.freebsd.org/src/commit/?id=b06338167d64 ("ROUTE_MPATH and FIB_ALGO") Fixes: https://cgit.freebsd.org/src/commit/?id=7e1affa242ca ("revise divert-to and divert-reply")
MFC after: 1 week
- s/pointr/pointer/ Obtained from: NetBSD MFC after: 3 days
Reviewed by: fuz Approved by: fuz (mentor) MFC after: 1 month Pull Request: https://github.com/freebsd/freebsd-src/pull/2352
One is the parameter name.
- s/uneccessarily/unnecessarily/ Obtained from: NetBSD MFC after: 3 days
- s/modifing/modifying/ MFC after: 3 days
- s/errornous/erroneous/ MFC after: 3 days
- s/predifined/predefined/ MFC after: 3 days
- s/modifing/modifying/ MFC after: 3 days
- s/modifing/modifying/ MFC after: 3 days
- s/varaiables/variables/ MFC after: 3 days
- s/staring/starting/ MFC after: 5 days
SpacemiT has only one T. Fixes: https://cgit.freebsd.org/src/commit/?id=dcb10e3add17 ("clk: Initial support for the SpacemiT K1 clock control units") Sponsored by: The FreeBSD Foundation
- Fix whitespace - Replace JH7110_GPIO_READ with RD4 (and WR4) - Trim headers - Explicit conditional checks - Use correct method typedefs MFC after: 3 days Sponsored by: The FreeBSD Foundation
Reviewed by: kib, markj Differential Revision: https://reviews.freebsd.org/D58858
Several lines in the flags-parsing block used spaces instead of tabs for indentation. Convert them to tabs to match style(9). No functional change. Sponsored by: Google LLC (GSoC 2026)
Use tabs consistently in #defines No functional change intended. Reported by: Hannes Elfert MFC after: 1 week MFC to: stable/15
Reviewed by: markj Differential Revision: https://reviews.freebsd.org/D59161
PR: https://bugs.freebsd.org/bugzilla/show_bug.cgi?id=298044 MFC after: 3 days
Fixes: https://cgit.freebsd.org/src/commit/?id=7f3b46fe54f1 ("ndp: Add support for Gratuitous Neighbor...")
Reported by: ngie@ Differential Revision: https://reviews.freebsd.org/D42156
Fixes: https://cgit.freebsd.org/src/commit/?id=a22fa5ec74e0 ("laoder.efi: Fix error in download protcol") Sponsored by: Netflix Differential Revision: https://reviews.freebsd.org/D59434
- s/ascic/ASCII/ MFC after: 3 days
- s/untill/until/ MFC after: 3 days
Reviewed by: olce Approved by: olce (mentor) MFC after: 2 weeks Sponsored by: FreeBSD Foundation Differential Revision: https://reviews.freebsd.org/D59400
Reviewed by: olce Approved by: olce (mentor) MFC after: 2 weeks Sponsored by: FreeBSD Foundation Differential Revision: https://reviews.freebsd.org/D59409
Fixes: https://cgit.freebsd.org/src/commit/?id=dcb10e3add17 Reported by: Bruno Banelli <bruno.banelli@sartura.hr> (p_idx typo)
In getcwd_logical(), test for a '.' or '..' component in one of the most
straightforward and intelligible ways possible.
In particular, this removes a superfluous re-test of the the component's
first character being '.' when the first one did not pass and, more
importantly, prevents the second test from relying on a side-effect in
the first.
While here, for better clarity, replace the loop that searches for '/'
with a simple call to strchrnul().
Add high-level comments about what is going on.
While here, test explicitly that pointed 'char' values are not 0 ('\0')
(style(9)).
While here, separate the successive steps of getcwd_logical() with blank
lines.
No functional change (intended).
Discussed with: emaste
Fixes: https://cgit.freebsd.org/src/commit/?id=2df923c5d2d0 ("pwd: Clean up and adopt POSIX semantics")
MFC after: 3 days
Sponsored by: The FreeBSD Foundation
Differential Revision: https://reviews.freebsd.org/D59709
Wrap long lines, related to the nullfs mounts over regular files and sockets type checks. Also fix indent. Sponsored by: The FreeBSD Foundation MFC after: 1 week
This automatically computes the correct PKG_CONFIG_PATH with LOCALBASE from the environment (when set) or from the "user.localbase" sysctl, in this order. Reviewed by: des Approved by: des Sponsored by: The FreeBSD Foundation Differential Revision: https://reviews.freebsd.org/D57246
Merge commit 'e988c5eab5231646c612d35ff5b16122cebfbf6a'
In delete(), when copying the deleted character to the d_char buffer, don't assume that it fits. utf8_prev() may return a sequence of more than 5 bytes. In insert_utf8(), fix the copy-up of the line. We extended the line by "len" bytes, so "temp" has to be repositioned accordingly. Compare with plain insert(). Use sizeof when copying to buffers instead of hard-coding buffer sizes. Don't dynamically allocate d_char, there is no need. Fixes: https://cgit.freebsd.org/src/commit/?id=62fba0054d9e ("ee: add unicode support") Reported by: Sayono Hiragi (overflow in delete()) Reviewed by: bapt MFC after: 2 weeks Sponsored by: The FreeBSD Foundation Differential Revision: https://reviews.freebsd.org/D57996
Otherwise one can't easily attach gdb to ee. Reviewed by: bapt MFC after: 2 weeks Sponsored by: The FreeBSD Foundation Differential Revision: https://reviews.freebsd.org/D57997
When looking up self process we can use `ps -p $$` directly rather than grep which may find other processes ending in the expected PID. Sponsored by: Dell Inc. Reviewed by: markj, vangyzen Differential Revision: https://reviews.freebsd.org/D58019
This is the last remaining piece of GPL software in the base system. The installer transitioned to bsddialog four years ago, and the last remaining dialog consumer, dpv, was turned off more than two years ago. Retire dpv, libdpv, libfigpar (used only by dpv), and dialog itself. Reviewed by: dteske Differential Revision: https://reviews.freebsd.org/D55424
Changes: https://github.com/eggert/tz/blob/2026c/NEWS MFC after: 3 days
MFC after: 1 week
Full release notes are available at https://www.openssh.com/txt/release-10.4 Selected highlights from the release notes: Potentially-incompatible changes -------------------------------- * sshd(8): configuration dump mode ("sshd -G") now writes directives in mixed case (e.g. "PubkeyAuthentication") whereas previously it emitted only lower-case names. * ssh(1), sshd(8): make the transport protocol stricter by disconnecting if the peer sends non-KEX messages during a post- authentication key re-exchange. Previously a malicious peer could continue sending non-key exchange messages without penalty. These would be buffered, causing memory to be wasted up until the connection terminated or the server/client hit a memory limit. Implementations that do not restrict messages sent during key exchange as per RFC4253 section 7.1 may be disconnected. Reported by Marko Jevtic. Changes since OpenSSH 10.3 ========================== This release contains a number of security fixes as well as general bugfixes and a couple of new features. Security ======== * sftp(1): when downloading files on the command-line using "sftp host:/path .", a malicious server could cause the file to be downloaded to an unexpected location. This issue was identified by the Swival Security Scanner. * scp(1): when copying files between two remote destinations, do not allow a malicious server to write files to the parent directory of the intended target directory. This issue was identified by the Swival Security Scanner. * sshd(8): DisableForwarding=yes didn't override PermitTunnel=yes as it was documented to do. Note that PermitTunnel is not enabled by default. Reported independently by Huzaifa Sidhpurwala of Redhat and Marko Jevtic. * sshd(8): avoid a potential pre-authentication denial of service when GSSAPIAuthentication was enabled (this feature is off by default). This was not mitigated by MaxAuthTries, but would be penalised by PerSourcePenalties. This was reported by Manfred Kaiser of the milCERT AT (Austrian Ministry of Defence). * sshd(8): fix a number of cases where the minimum authentication delay was not being enforced. Reported by the Orange Cyberdefense Vulnerability Team. * ssh(1): fix a possible client-side use-after-free if the server changes its host key during a key reexchange. This was reported by Zhenpeng (Leo) Lin of Depthfirst. New features ------------ * All: add experimental support for a composite post-quantum signature scheme that combines ML-DSA 44 and Ed25519 as specified in draft-miller-sshm-mldsa44-ed25519-composite-sigs. This scheme is not enabled by default. To use it, you'll need to add it to HostKeyAlgorithms, PubkeyAcceptedAlgorithms, etc. Keys may be generated using "ssh-keygen -t mldsa44-ed25519". Bugfixes -------- * sshd(8): avoid sending observably different messages for valid vs invalid users in GSSAPIAuthentication (disabled by default). * ssh(1), sshd(8): fix several bugs that incorrectly classified bulk traffic as interactive. bz3972, bz3958 * ssh-keygen(1), ssh-add(1): skip unsupported key types when downloading resident keys from a FIDO token. Previously, downloads would abort when one was encountered. GHPR657 Sponsored by: The FreeBSD Foundation Differential Revision: https://reviews.freebsd.org/D58083
We don't use this note type today, but as a general purpose ELF diagnostic tool readelf(1) ought to decode it. References: https://fedoraproject.org/wiki/Changes/Package_information_on_ELF_objects https://systemd.io/ELF_PACKAGE_METADATA/ Reviewed by: fuz Sponsored by: The FreeBSD Foundation Differential Revision: https://reviews.freebsd.org/D47524
Differential Revision: https://reviews.freebsd.org/D58333 Approved by: ivy MFC after: 3 days Changelog: https://github.com/vstakhov/libucl/releases/tag/0.9.4
Add -P as shorthand for --pause-before-cleanup. MFC after: 1 week Reviewed by: ngie Differential Revision: https://reviews.freebsd.org/D56613
Release notes at
https://community.nlnetlabs.nl/t/unbound-1-25-2-released
Merge commit 'c68e7bcd81d62e9f5364c6da22fd9917976acf85'
Security: CVE-2026-14586
Security: CVE-2026-32665
Security: CVE-2026-40691
Security: CVE-2026-41637
Security: CVE-2026-42955
Security: CVE-2026-44621
Security: CVE-2026-44687
Security: CVE-2026-44690
Security: CVE-2026-46582
Security: CVE-2026-50045
Security: CVE-2026-50046
Security: CVE-2026-50243
Security: CVE-2026-50248
Security: CVE-2026-50251
Security: CVE-2026-50252
Security: CVE-2026-52863
Security: CVE-2026-54478
Security: CVE-2026-55708
Security: CVE-2026-55717
Security: CVE-2026-55973
Security: CVE-2026-55990
Security: CVE-2026-55991
Security: CVE-2026-56416
Security: CVE-2026-56444
Update the mt76/zzz_fw_ports_fwget.sh script to set fwget to download mt7921 and mt7925 rather than the these days non-existent mt792x flavor. Sponsored by: The FreeBSD Foundation MFC after: 30 days Differential Revision: https://reviews.freebsd.org/D57242
Reviewed by: markj, ngie Sponsored by: The FreeBSD Foundation MFC after: 1 week Differential revision: https://reviews.freebsd.org/D58458
[NFC][ELF][PPC64] Pass address not offset to writePPC64LoadAndBranch (#212275) Every caller currently subtracts the TOC base in its argument, so move that into common code inside writePPC64LoadAndBranch. This will also allow a different computation to be used in some cases in a future commit. Note that offset is now unsigned not signed; even previously, all arguments were uint64_t, and all uses are unsigned, so making it signed doesn't make much sense. MFC after: 1 week
[ELF][PowerPC] Don't assume TOC pointer is valid in IPLT entries (#207555) Unlike normal PLT entries, IPLT entries can be called indirectly even when in PIEs/DSOs, and so there's no guarantee on what's in the TOC pointer register at that time. Therefore we must emit variants of the existing code that work without it, whether r12-relative (playing the same role as MIPS's $25) in the same number of instructions, or first retrieving PC in an i386-like manner, being careful not to clobber LR. On 32-bit PowerPC even direct calls to IPLT entries face the same issue, since we'd use the TOC base of the resolver, which may not be the same as the caller, even within the same object. Normal canonical PLTs still look broken on 64-bit PowerPC as they use the TOC pointer register too, and similarly on 32-bit PowerPC for PIEs. We should probably treat these cases the same as PIE on i386 (except including PDEs for 64-bit PowerPC), where it's an error due to the use of %ebx in PLT entries. Bump LLD_FREEBSD_VERSION for this fix as otherwise an existing system linker will be deemed new enough to use and produce broken kernels for TARGET=powerpc (regardless of TARGET_ARCH/MACHINE/MACHINE_ARCH) builds. PR: https://bugs.freebsd.org/bugzilla/show_bug.cgi?id=294369 MFC after: 1 week
[libunwind][PPC64] Fix unw_getcontext corrupting callee-saved VSX registers on LE (#198371) This is the first of two independent fixes for libunwind on ppc64le (ELFv2 ABI, little-endian), where two separate bugs together cause SIGSEGV during backtracing. This commit addresses the VSX register corruption; the TOC-restore fault is handled in a follow-up. Both were discovered while debugging lang/rust build failures with RUST_BACKTRACE=1 on FreeBSD/powerpc64le (IBM POWER9). On ppc64le, `unw_getcontext` saves each VS register with an in-place `xxswapd n, n` followed by `stxvd2x`. The swap is needed because `stxvd2x` stores doublewords in the wrong order on LE. However, the macro never applies a second `xxswapd` to restore the register after the store, so all 64 VS registers are permanently corrupted on return from `unw_getcontext`. This affects every callee-saved VSX register: f14-f31 (VSR14-VSR31) and VR20-VR31 (VSR52-VSR63). After `_Unwind_Backtrace` returns, any code that uses these registers sees wrong values. In practice this manifests as SIGSEGV inside hashbrown's `reserve_rehash`: VR20-VR31 are corrupted before a SIMD comparison loop runs, producing an out-of-bounds access. Fix: add a second `xxswapd n, n` after the `stxvd2x` store. Since `xxswapd` is its own inverse, the pair is a no-op on the architectural register while still writing the correctly byte-swapped value to memory. MFC after: 1 week
libarchive 3.8.9 ChangeLog: https://github.com/libarchive/libarchive/compare/v3.8.7...v3.8.9 Obtained from: libarchive Vendor commit: 27cbc7827172698143e440801fc0ba39ccb4f1f5 MFC after: 2 weeks
On some platforms, e.g. Linux Clang 22.1.8 / glibc 2.43, strchr() now implements the C23 behaviour where passing a const pointer to strchr() also returns a const pointer. This breaks libucl during the bootstrap build, since it assumes the return value is always a mutable pointer. Instead of assigning directly to params->prefix (which is const), use a non-const temporary variable and assign the result after we've done the modification. MFC after: 1 week Reviewed by: bofh, bapt Sponsored by: The FreeBSD Foundation Differential Revision: https://reviews.freebsd.org/D58490
On some platforms, e.g. Linux Clang 22.1.8 / glibc 2.43, strchr() now implements the C23 behaviour where passing a const pointer to strchr() also returns a const pointer. This breaks mandoc during the bootstrap build, since it assumes the return value is always a mutable pointer. In read.c, make the existing temporary pointer const, and for the mandoc_asprintf() call, add a new mutable local. In mdoc.c and out.c, since the data is mutable and is mutated here, remove const from the temporary pointers. MFC after: 1 week Reviewed by: fuz Sponsored by: The FreeBSD Foundation Differential Revision: https://reviews.freebsd.org/D58495
On some platforms, e.g. Linux Clang 22.1.8 / glibc 2.43, strchr() now implements the C23 behaviour where passing a const pointer to strchr() also returns a const pointer. This breaks libelftc during the bootstrap build, since it assumes the return value is always a mutable pointer. Since the returned pointer is never modified in either case, make it const. MFC after: 1 week Reviewed by: jkoshy, markj, dim, emaste Sponsored by: The FreeBSD Foundation Differential Revision: https://reviews.freebsd.org/D58497
Release notes at
https://community.nlnetlabs.nl/t/unbound-1-26-0-released
Merge commit '84ffc29dc8ddb0c946db5cb3b3c1310bec6a9e6c'
Reviewed by: markj MFC after: 1 week Differential Revision: https://reviews.freebsd.org/D58634
Found with: clang -Werror=assign-enum
Changes: https://github.com/libexpat/libexpat/blob/R_2_8_3/expat/Changes Security: CVE-2026-72522 MFC after: 1 week
MFC after: 3 days
Release notes at https://www.sqlite.org/releaselog/3_53_3.html. Obtained from: https://www.sqlite.org/2026/sqlite-autoconf-3530300.tar.gz MFC after: 2 weeks Merge commit 'e698feec080925c6cffa9ec31be884daa5cea536'
This change syncs the lib/libc/c063 NetBSD tests with FreeBSD. This does two things: - Addresses bogus tautologically true assertions flagged by clang and gcc with ATF 0.22+ [1]. - Brings in some new test coverage. Obtained from: NetBSD (date tag: `20260818UTC`) MFC after: 2 weeks 1. https://github.com/freebsd/atf/pull/72
This is a rollup commit from upstream to fix: Handle signature_algorithms_cert extension in key-only context Avoid double free of qrx in port_default_packet_handler() Avoid full read buffer allocation when buffering DTLS next-epoch records ssl/record/methods/dtls_meth.c: lower the unprocessed_rcds queue limit ssl/record: remove dead DTLS processed_rcds record queue Fix heap buffer overflow (8-byte OOB write) in AES-WRAP-PAD unwrap CMP unexpected sender DN used as format string in ERR_raise_data() Add test for CVE-2026-63073 Add a test for restricting growth in cmp cert cache Fix unbounded cert cache growth in cmp Don't store ACK-only frames in TX history for QUIC. Add test for CVE-2026-63076 Fix Remote NULL deref in ossl_cmp_calc_protection() via crafted protectionAlg Approved by: so Obtained from: OpenSSL Security: FreeBSD-SA-26:61.openssl Security: CVE-2026-14457 Security: CVE-2026-18798 Security: CVE-2026-54874 Security: CVE-2026-63072 Security: CVE-2026-63073 Security: CVE-2026-63074 Security: CVE-2026-63076
Suggested by: cy, emaste Reviewed by: cy Differential Revision: https://reviews.freebsd.org/D55929
This is a security bugfix release. Please see the related merge commit
for more details.
Maintainer note: `quic_ackm.h`'s conflict was resolved by taking
the upstream version of the file verbatim.
Conflicts:
crypto/openssl/include/internal/quic_ackm.h
MFC after: 3 days
Merge commit '248da023ae5ea7292930ac5d715d88b87e2e6f46'
crypto/openssl: update generated content to match 3.5.8 release This contains 2 new manpages as well as some minor manpage content changes. MFC with: 78e936b2d
crypto/openssl: add manpages missed in related commit MFC with: 78e936b2d Fixes: https://cgit.freebsd.org/src/commit/?id=0d4d0f3a9 ("crypto/openssl: update generated content ...") Reported by: Jenkins CI
More information about what's included in the new release can be found [here](https://github.com/freebsd/lutok/compare/lutok-0.4...lutok-0.6.2). MFC after: 1 week Merge commit '447d4fe61d8bf55ffa21698ecee0ac51010a9a0f'
The extra DELAY seems to no longer be needed and the dump_stack() is definitively a problem now. Remove all this. Sponsored by: The FreeBSD Foundation MFC after: 3 days
Changes: https://github.com/libexpat/libexpat/blob/R_2_8_4/expat/Changes Security: CVE-2026-66046 Security: CVE-2026-76641 Security: CVE-2026-76957 Security: CVE-2026-76956 MFC after: 1 week
MFV: file 5.48 MFC after: 1 week
libmagic: Add swap.c and magic.h to SRCS. file 5.48 added swap.c and swap.h for byte-swapping operations, which are required on hosts that lack <byteswap.h> or <sys/bswap.h> (e.g. macOS cross-building or older FreeBSD bootstrap environments). Also add magic.h to SRCS so object files depend on the generated header before compiling, avoiding falling back to the host's /usr/include/magic.h during parallel builds. Reported by: wosch MFC after: 1 week Fixes: https://cgit.freebsd.org/src/commit/?id=7af41682a96b ("MFV: file 5.48")
This fixes the build with gcc 14:
/usr/src/contrib/kyua/engine/prepare/prepare_all.cpp:56:16: error: declaration of 'handler' shadows a member of 'engine::prepare::prepare_all' [-Werror=shadow]
56 | for (auto& handler : prepare::handlers()) {
| ^~~~~~~
In file included from /usr/src/contrib/kyua/engine/prepare/prepare_all.hpp:35,
from /usr/src/contrib/kyua/engine/prepare/prepare_all.cpp:29:
/usr/src/contrib/kyua/engine/prepare/prepare.hpp:51:15: note: shadowed declaration is here
51 | class handler {
| ^
Fixes: https://cgit.freebsd.org/src/commit/?id=edb230c4af499203d7a6894b3711fe6574b26040
Reviewed by: igoro, rlibby, ngie
MFC after: 3 days
Sponsored by: The FreeBSD Foundation
Differential Revision: https://reviews.freebsd.org/D59346
These files provide no value in the FreeBSD tree proper and change frequently, depending on what machine I generate the release tarball on (and what versions of autotools are on the host). Nuke the autogenerated files to avoid bloating commit history/the tree. MFC after: 3 days Requested by: Benjamin Jacobs <freebsd@dev.thsi.be>
This fixes the build with gcc 16 after the changes to -Wunused*[0]. [0] https://gcc.gnu.org/gcc-16/porting_to.html#changes-to-wunused Reviewed by: markj MFC after: 3 days Sponsored by: The FreeBSD Foundation Differential Revision: https://reviews.freebsd.org/D59537
pkgbase splits programs and their tests into separate packages (e.g. unifdef/yacc/indent/file2c -> toolchain, jail -> jail, etc), while their tests always land in the generic tests package. Installing just the tests package therefore leaves these suites unable to find the binary they exercise, and they fail confusingly instead of skipping. Rather than force all packages to be installed for the tests package, make the tests cope with the missing packages. Add require.progs (or, for TAP/PLAIN-style suites with no atf_set hook: TEST_METADATA required_programs) for the binary under test in: bectl, certctl, ctfconvert, dhclient, file2c, indent, ipfw, jail, lastcomm, mixer, newsyslog, nvmecontrol, pfctl, praudit, sa, syslogd, tunefs, unifdef, yacc, ztest. I skipped adding data for all the binaries in the base runtime package for simplicity. Assisted-by: Claude Code (Sonnet 5) Sponsored by: Netflix Reviewed by: ngie Differential Revision: https://reviews.freebsd.org/D58262
Changes: https://github.com/eggert/tz/blob/2026d/NEWS MFC after: 3 days
tzcode: Update to 2026c MFC after: 1 week
libc/stdtime: Catch up with tzcode 2026d Fixes: https://cgit.freebsd.org/src/commit/?id=212e35249432 ("tzcode: Update to 2026c")
MFC after: 3 days
The VF-up script repeatedly deletes the first IPv4 address from hn until ifconfig fails. Netlink-based ifconfig returns success when no address remains, so the script can loop forever and block subsequent devd work. Use ipv4_down() from the already-sourced network.subr to enumerate the addresses and delete each explicitly. This also handles multiple addresses without relying on the exit status of an empty deletion. The behavior depends on ifconfig, not the VF hardware, so no device-specific fallback is needed. Fixes: https://cgit.freebsd.org/src/commit/?id=c68595695679 ("hyperv: Add VF bringup scripts and devd rules.") MFC after: 2 weeks Sponsored by: BBOX.io
The default VF-up script moves IPv4 configuration from the synthetic interface to a failover lagg, but leaves IPv6 unconfigured on the lagg. Adding the synthetic interface as a member can remove its IPv6 addresses, so deleting the remaining member addresses is not sufficient. Stop accepting router advertisements on the synthetic interface before adding it to the lagg. Leave its link-local address and IPv6 enable state in place for lagg's address, prefix and default-router cleanup, then disable IPv6 and explicitly remove any remaining addresses. This also covers members without a link-local address or with IPv6 already disabled. Replay the synthetic interface's IPv6 rc.conf configuration on the lagg, including aliases, prefix-derived addresses and legacy configuration names. Remap the variables in a subshell and reuse network.subr's IPv6 helpers. For SLAAC, solicit fresh router advertisements rather than copying learned addresses as permanent ones. Configure IPv6 independently of the IPv4 DHCP/static choice, and leave existing laggs alone on repeated invocation. Custom DHCPv6 clients and interface-scoped static routes still require the per-interface setup hook. Fixes: https://cgit.freebsd.org/src/commit/?id=c68595695679 ("hyperv: Add VF bringup scripts and devd rules.") MFC after: 2 weeks Sponsored by: BBOX.io
Sponsored by: Chelsio Communications
This change allows libcxgb4 to post send/recv drain work requests after a QP is flushed and schedules a CQE on the software CQ for the application to note the drain's completion. Sponsored by: Chelsio Communications
Remove it because it is not needed. If we're flushing the qp, it is because we've had a fatal error, so just set the user qp state to ERR and continue on. Sponsored by: Chelsio Communications
[PATCH 02/31] FreeBSD OFED support for DPDK MLX5 PMD Extend mlx5dv_get_qp() to return UAR mapping info. This can allow a process to share its doorbell access with secondary process by re-mmap the UAR address on the device and make it accessible as a user space address. Differential revision: https://reviews.freebsd.org/D32176 MFC after: 1 month
Import rdma-core upstream commit d2389b34ccc5 ("mlx5: Add missing include
file in mlx5dv.h").
Including <sys/types.h> keeps mlx5dv.h self-contained: off_t, used by the
new uar_mmap_offset field, is now always defined. The man page gains the
uar_mmap_offset line upstream documents in 5dd3ead7b5f5.
Sponsored by: NVidia networking
MFC after: 1 month
PATCH 16/31] FreeBSD OFED support for DPDK MLX5 PMD a) verbs: Annotate ibv_wc helpers with endian This follows the scheme as used in the wc by introducing a ibv_wc_read_invalidated_rkey to access the host endian invalidated_rkey value with proper annotations. This is just an inline wrapper to allow sparse to work sensibly, not really a good reason to add another driver entry point. b) verbs: Annotate ibv_send_wr with endian This follows the scheme as used in the wc by introducing a transparent union in ibv_send_wr to indicate the invalidate_rkey is in host endian. c) mlx5: Add sparse annotations. Differential revision: https://reviews.freebsd.org/D32190 MFC after: 1 month
Fix is based on upstream rdma-core commit 14a0fc824f16 ("rdma-core/irdma:
Implement device supported verb APIs").
Use ib_wr->invalidate_rkey instead of ib_wr->imm_data for
IBV_WR_SEND_WITH_INV and IBV_WR_LOCAL_INV. The two share a union, but
imm_data is __be32 while the rkey to invalidate is host order, so with the
new endian annotations this path reads the wrong member. Upstream irdma
has always used invalidate_rkey here.
Sponsored by: NVidia networking
MFC after: 1 month
[PATCH 18/31] FreeBSD OFED support for DPDK MLX5 PMD Provide two implementations of mlx5dv_init_obj, one that has the historical behaviour that has existed until now of returning the void **uar and a new version that returns the 'void *' version renamed to cq_uar. Differential revision: https://reviews.freebsd.org/D32192 MFC after: 1 month
Actually provide the two ABI versions this commit describes: export mlx5dv_init_obj@@MLX5_1.2 with the current behaviour and a compat mlx5dv_init_obj@MLX5_1.0 that restores the historical 'void **' value, and list the symbol under both versions in libmlx5.map. Binaries linked against the old version keep resolving. Sponsored by: NVidia networking MFC after: 1 month
[PATCH 19/31] FreeBSD OFED support for DPDK MLX5 PMD Add a new DV API mlx5dv_set_context_attr() to enable setting an external memory allocator. This API will allow the application to use specific decisions about the memory allocation of HW resources (e.g. DV objects). Some examples are managing numa pinning per object, managing a hugepages resource pool, shared memory regions. Also extend mlx5dv_get_qp() to return UAR mapping info. This can allow a process to share its doorbell access with secondary process by re-mmap the UAR address on the device and make it accessible as a user space address. Differential revision: https://reviews.freebsd.org/D32193 MFC after: 1 month
Match a doorbell to its page with a range check instead of a page-boundary mask, so mlx5_free_db() also works for external allocations that are not page aligned and no longer leaks the doorbell page. Sponsored by: NVidia networking MFC after: 1 month
PATCH 20/31] FreeBSD OFED support for DPDK MLX5 PMD a) Use flag MLX5DV_CONTEXT_FLAGS_MPW_ALLOWED to indicate hardware supports multi packet WQE and it's enabled in SQ context. Flag MLX5DV_CONTEXT_FLAGS_MPW is deprecated, shall not be used in new applications. b) Report if enhanced multi packet send WQE is supported through mlx5 direct verbs. Differential revision: https://reviews.freebsd.org/D32194 MFC after: 1 month
Interpret the kernel's raw multi_pkt_send_wqe value (0, 1 or 3) instead of masking it against a bitmask the kernel never emits, which silently disabled MPW. Sponsored by: NVidia networking MFC after: 1 month
[PATCH 23/31] FreeBSD OFED support for DPDK MLX5 PMD a) Expose tag matching capabilities and show them in ibv_devinfo. b) verbs: Introduce tag matching SRQ Introducing tag matching SRQ (TM-SRQ), which retains basic semantic of regular SRQ, reports completions to own CQ, and has additional tag based message receiving mechanism. Detailed description for the TM-SRQ usage was added into Documentation/tag_matching.md c) mlx5: Add support to tag matching SRQ type Create command QP for a tag-matching SRQ. Command QP used for inserting/removing entries from the tag matching list. This command QP is hidden from the users in the mlx5_srq structure. New verb ibv_post_srq_ops() will be added in next patch to use it. Differential revision: https://reviews.freebsd.org/D32198 MFC after: 1 month
Reject TM-SRQ creation when tm_cap.max_ops is 0, so the internal command QP's send queue is not sized to zero work-queue entries. Sponsored by: NVidia networking MFC after: 1 month
[PATCH 24/31] FreeBSD OFED support for DPDK MLX5 PMD Software parsing (SWP) is a feature that can be used to instruct the device to stop using its internal parser and to parse packets on the transmit path according to offsets set for each packet. Through this feature, the device allows the handling of checksum and LSO by the hardware according to the location of IP and TCP/UDP headers. Report various SW parsing capabilities and supported QP types through mlx5 direct verbs interface. Differential revision: https://reviews.freebsd.org/D32199 MFC after: 1 month
[PATCH 26/31] FreeBSD OFED support for DPDK MLX5 PMD A Multi-Packet RQ is a receive queue where multiple packets are written to the same WQE. Each message starts in the beginning of a stride. The total size of the scatter elements of each WQE is determined upon RQ creation and all the posted WQEs should meet the determined size. A Multi-Packet RQ reduces the number of needed post-recv operations thus increasing performance. It reduces memory footprint by allowing each packet to consume a different number of strides instead of the whole WR. Differential revision: https://reviews.freebsd.org/D32201 MFC after: 1 month
[PATCH 27/31] FreeBSD OFED support for DPDK MLX5 PMD
a) Add needed definitions to allow creation of a Multi-Packet RQ
using the mlx5 direct verbs interface.
In order to create a Multi-Packet RQ, one needs to provide a
mlx5dv_wq_init_attr containing the following information in its
striding_rq_attrs struct:
- single_stride_log_num_of_bytes: log of size of each stride
- single_wqe_log_num_of_strides: log of number of strides per WQE
- two_byte_shift_en: When enabled, hardware pads 2 bytes of zeros
before writing the message to memory (e.g. for IP alignment).
b) Add a helper function to verify 64 bit comp mask
The common check for a mask is as follows:
if (comp_mask & ~COMP_MASK_SUPPORTED_VALUES)
return EINVAL;
This can cause an issue when using 64 bit mask if the supported variable
is signed 32 bit: It will be bitwise inverted and then zeroed to 64
bits. Hence, wrong bits in the mask that exceed 32 bits will not raise
an error but would be ignored.
To fix this, a helper function is added in driver.h to be used by providers
code and fix wrong mask checks where the above was found to be applicable.
Differential revision: https://reviews.freebsd.org/D32202
MFC after: 1 month
[PATCH 28/31] FreeBSD OFED support for DPDK MLX5 PMD a) In order to enable offloading such as checksum and LRO for incoming tunneling traffic, the QP should be created with tunnel offloads flag - MLX5DV_QP_CREATE_TUNNEL_OFFLOAD. b) Reports capability of which tunneling type supports the tunneling offloads. c) verbs: Add support in RSS of the inner packet Some user space application would like to do RSS on the inner packet fields instead of the outer. When user will set the IBV_RX_HASH_INNER bit with one of the other hash fields, then the RSS will be on the inner packet. Differential revision: https://reviews.freebsd.org/D32203 MFC after: 1 month
Fix is based on upstream rdma-core commit ee54f9d5348f ("mlx5: Add
loopback flags to QP creation").
Validate create_flags against a supported-flags mask instead of rejecting
anything that is not MLX5DV_QP_CREATE_TUNNEL_OFFLOADS. A create_flags of
0 is now accepted, unknown bits are rejected, and the vendor flag is OR-ed
in rather than assigned. The loopback flags that commit also adds do not
exist in this tree, so only the validation is taken.
Sponsored by: NVidia networking
MFC after: 1 month
[PATCH 30/31] FreeBSD OFED support for DPDK MLX5 PMD a) Allow verbs applications packet steering of GRE tunneled traffic. Adding GRE flow specification based on RFC 2890. GRE consists of flags, protocol and key fields. IPv4 protocol 47 (IPPROTO_GRE) can be used when GRE packets are encapsulated in IPv4. b) verbs: Add MPLS flow specification filter Add MPLS flow specification based on RFC 3032. MPLS spec defined with label field which includes the label value and additional parameters such as: BoS, TC and TTL. MPLS allows stacking multiple labels in sequence. In addition, the MPLS header can be encapsulated on top of different layers, e.g.: ETH, IP (rfc4023), UDP (rfc7510), GRE (rfc4023). Therefore, when using the flow creation verb, the application should organize the spec filters list in the command in an ordered manner, such that reflects the actual protocol stack of the packet, to determine the exact position of the MPLS headers in the protocol stack. c) mlx5: Report MPLS tunnel offload capabilities through mlx5 direct verbs This patch exposes the mlx5 device's capability to offload MPLS based tunnel protocols via DV API. The possible protocols are: - Control word + MPLS over GRE. - Control word + MPLS over UDP. Differential revision: https://reviews.freebsd.org/D32205 MFC after: 1 month
Import rdma-core upstream commit ff01da2c5ac2 ("verbs: Fix typo in copying
IBV_FLOW_SPEC_UDP/TCP 'val'").
The TCP/UDP filter value was copied using the size of the IPv4 filter,
overrunning the destination subobject.
Also accept inner MPLS flow specs, matching 6d6f29721eea ("verbs: Allow
creation of inner MPLS flow spec"). Upstream deliberately has no inner
case for GRE - it is a tunnel header, matched only in the outer stack - so
that part of the review comment is not applied.
Sponsored by: NVidia networking
MFC after: 1 month
Release notes at
https://community.nlnetlabs.nl/t/unbound-1-26-1-released
Merge commit '120aa088f4807126af42a652a07c5090e7294fcf'
Security: CVE-2026-77860
Security: CVE-2026-77955
Security: CVE-2026-78227
Security: CVE-2026-80225
Security: CVE-2026-81634
Security: CVE-2026-81642
Security: CVE-2026-82717
Security: CVE-2026-82720
Security: CVE-2026-85501
Fix -Wformat diagnostic after #190965 (#193704)
Fixes libunwind compiler diagnostic when building with clang after
034d4dcad6396d1241e8262e69871b8d61da7e4f:
```
In file included from libunwind/src/libunwind.cpp:31:
In file included from libunwind/src/UnwindCursor.hpp:52:
libunwind/src/CompactUnwinder.hpp:339:46: error: format specifies type 'unsigned long long' but the argument has type 'uint64_t' (aka 'unsigned long') [-Werror,-Wformat]
338 | "function starting at 0x%llX",
| ~~~~
| %lX
339 | compactEncoding, functionStart);
| ^~~~~~~~~~~~~
libunwind/src/config.h:215:63: note: expanded from macro '_LIBUNWIND_DEBUG_LOG'
215 | #define _LIBUNWIND_DEBUG_LOG(msg, ...) _LIBUNWIND_LOG(msg, __VA_ARGS__)
| ~~~ ^~~~~~~~~~~
libunwind/src/config.h:181:45: note: expanded from macro '_LIBUNWIND_LOG'
181 | fprintf(stderr, "libunwind: " msg "\n", __VA_ARGS__); \
| ~~~ ^~~~~~~~~~~
In file included from libunwind/src/libunwind.cpp:31:
In file included from libunwind/src/UnwindCursor.hpp:52:
libunwind/src/CompactUnwinder.hpp:458:39: error: format specifies type 'unsigned long long' but the argument has type 'uint64_t' (aka 'unsigned long') [-Werror,-Wformat]
457 | "function starting at 0x%llX",
| ~~~~
| %lX
458 | encoding, functionStart);
| ^~~~~~~~~~~~~
libunwind/src/config.h:215:63: note: expanded from macro '_LIBUNWIND_DEBUG_LOG'
215 | #define _LIBUNWIND_DEBUG_LOG(msg, ...) _LIBUNWIND_LOG(msg, __VA_ARGS__)
| ~~~ ^~~~~~~~~~~
libunwind/src/config.h:181:45: note: expanded from macro '_LIBUNWIND_LOG'
181 | fprintf(stderr, "libunwind: " msg "\n", __VA_ARGS__); \
| ~~~ ^~~~~~~~~~~
```
This avoids a -Werror failure with clang >= 23.
MFC after: 3 days
MFC after: 1 week
Improve manpage compatability by dropping an abbreviation macro that compiler upstreams do not want, causing manuals using it to render incorrectly on other toolchains or operating systems. It was only used in three manuals, which are amended in this commit. Reviewed by: fuz, markj MFC after: no, these did not MFC Requested by: Ingo Schwarze <schwarze@openbsd.org> Fixes: https://cgit.freebsd.org/src/commit/?id=db3884b03989 ("contrib/mandoc: add -ieee754-2008") Fixes: https://cgit.freebsd.org/src/commit/?id=63cd0841de76 ("document .St -ieee754-2008 in mdoc") Differential Revision: https://reviews.freebsd.org/D59880
In the event that NDEBUG is specified in CFLAGS--which is most likely triggered via `MK_ASSERT_DEBUG` == "no" -- all assert(3) statements are optimized out by design. This breaks the assert(3) tests as they specifically rely on assert(3) actually raising a `SIGABRT` instead of quietly succeeding. Skip both tests if `NDEBUG` is set. There's no sense running the `assert(true)` test if the result could instead be a false positive. MFC after: 2 weeks
contrib/atf: upgrade to 0.26
This new version gets contrib/atf in line with the published version of
the package up on [GitHub][github].
Please note that this release is a major update: it contains several
bugfixes and feature enhancements. Most were already present in the
FreeBSD src repo, but there were a large number of changes made in
addition to what was already present.
MFC after: 1 month
Merge commit '9c78088348bcf1059d82a62bb57a0746309d4312'
Conflicts:
contrib/atf/.cirrus.yml
contrib/atf/Makefile.am
contrib/atf/Makefile.in
contrib/atf/aclocal.m4
contrib/atf/admin/ar-lib
contrib/atf/admin/check-style-common.awk
contrib/atf/admin/check-style.sh
contrib/atf/admin/compile
contrib/atf/admin/config.guess
contrib/atf/admin/config.sub
contrib/atf/admin/depcomp
contrib/atf/admin/install-sh
contrib/atf/admin/ltmain.sh
contrib/atf/admin/missing
contrib/atf/atf-c++/Makefile.am.inc
contrib/atf/atf-c++/atf-c++.3
contrib/atf/atf-c++/atf-c++.pc.in
contrib/atf/atf-c++/check.cpp
contrib/atf/atf-c++/check.hpp
contrib/atf/atf-c++/detail/Makefile.am.inc
contrib/atf/atf-c++/detail/process_test.cpp
contrib/atf/atf-c/Makefile.am.inc
contrib/atf/atf-c/atf-c.3
contrib/atf/atf-c/atf-c.pc.in
contrib/atf/atf-c/detail/Makefile.am.inc
contrib/atf/atf-c/detail/fs.c
contrib/atf/atf-c/detail/fs_test.c
contrib/atf/atf-c/detail/process_test.c
contrib/atf/atf-c/tc.c
contrib/atf/atf-c/tc.h
contrib/atf/atf-sh/Makefile.am.inc
contrib/atf/atf-sh/atf-check.cpp
contrib/atf/atf-sh/atf-check_test.sh
contrib/atf/bootstrap/Makefile.am.inc
contrib/atf/bootstrap/package.m4
contrib/atf/bootstrap/testsuite
contrib/atf/config.h
contrib/atf/configure
contrib/atf/configure.ac
contrib/atf/doc/Makefile.am.inc
contrib/atf/doc/atf-test-case.7
contrib/atf/m4/developer-mode.m4
contrib/atf/m4/libtool.m4
contrib/atf/m4/ltoptions.m4
contrib/atf/m4/ltsugar.m4
contrib/atf/m4/ltversion.m4
contrib/atf/m4/lt~obsolete.m4
contrib/atf/m4/module-defs.m4
contrib/atf/test-programs/Makefile.am.inc
[github]: https://github.com/freebsd/atf
atf-sh: install kmod_test added as part of the 0.26 release Fixes: https://cgit.freebsd.org/src/commit/?id=952231086100 ("contrib/atf: upgrade to 0.26")
integration_test: backport fix from freebsd/atf See the related merge commit for more details. MFC after: 28 days Obtained from: [freebsd/atf@2690b04c][upstream] Fixes: https://cgit.freebsd.org/src/commit/?id=95223108 ("contrib/atf: upgrade to 0.26") Merge commit '8fd09407c8e64fd26d095c9b00990c545577db5a' [upstream]: https://github.com/freebsd/atf/commit/2690b04cbc608af165a58426b7275c063b779f03
Changes: https://github.com/libexpat/libexpat/blob/R_2_8_5/expat/Changes Security: CVE-2026-93990 MFC after: 3 days
Verbatim import of these files, with unix line endings. Sponsored by: Netflix
Sponsored by: Netflix
This is a commit from upstream to fix: dtls: reset init_off before retransmitting a message Approved by: so Security: FreeBSD-SA-26:68.openssl Security: CVE-2026-84782
Changes: https://github.com/eggert/tz/blob/2026e/NEWS Briefly: Manitoba moves to permanent -05 on 2026-10-31. MFC after: 3 days
MFC after: 1 week
[Clang][ExprConst] Drop PRValue for nothrow new (#226753) A user defined operator new can accept prvalue for nothrow. However, it should not be a ConstExpr. Early returns by isUsableAsGlobalAllocationFunctionInConstantEvaluation instead of doing LValue evaluation. Also, move CheckPlacement new logic into new OpCode. This decouples checking from Interp.cpp to Compiler.cpp. This fixes "Assertion failed: (E->isGLValue() || E->getType()->isFunctionType() || E->getType()->isVoidType() || isa<ObjCSelectorExpr>(E->IgnoreParens())), function EvaluateLValue" when building the databases/mariadb123-server port. MFC after: 1 week
This is a security fix release addressing CVE High issues. Users are strongly encouraged to update to this version. See the release notes for the release for more details on what is being fixed. MFC after: 1 day Merge commit 'af5a659dc1cd2b0a6994f2cee1956d0bf50bb1a2'
A new manpage has been added and some source files have been refactored slightly, but by and large this is just a standard "version bump" update (3.5.8 -> 3.5.9). MFC after: 1 day MFC with: b3a31d78
zfs: Wire sha512 offload to the build FreeBSD main just got the CPUID_STDEXT4_SHA512 define. OpenZFS PR #18732
Revert "zfs: Wire sha512 offload to the build" This reverts commit cd61eb4f6681b13d98b6a7be252500ad30f05f74. Some people report module load failure due to undefined symbol. I don't have those problems myself, so it might be a question of full rebuild. But I don't have time right now, so just revert.
zfs: add btree.c to libzfs Fixes breakage from 34f9f5680 (openzfs/zfs@1f380a4f3)
Revert "zfs: add btree.c to libzfs" This reverts commit b4fdb5c236d00825c8e3ea8ffc7bc25dafb6abcf.
This reverts commit 3e3fd1fde8e168910edc538966111c0b5f03cd5f. This appears to break chainbooting with boot1.efi and similar scenarios with Root-on-ZFS scenarios. Revert until it's better understood. PR: https://bugs.freebsd.org/bugzilla/show_bug.cgi?id=296309 Sponsored by: Netflix Differential Revision: https://reviews.freebsd.org/D58071
This reverts commit 74654ba3b1b3bcf6ba8870a54310accbb6adbf0b. Apparently it breaks cross building from Linux for some reason. I'll admit I didn't even know we supported cross building from Linux.
pkg: Add -j and -r options This allows pkg(7) to be used to bootstrap a jail or chroot, and to recognize the -j and -r options and pass them through to pkg(8) if already bootstrapped. Note that this does not address the issue of repository keys. If using a signed package repository, you will still need to copy /usr/share/keys into the target environment before or after bootstrapping, or pkg will be unable to verify package signatures. MFC after: 1 week Reviewed by: imp, bapt Differential Revision: https://reviews.freebsd.org/D58165
Revert "pkg: Add -j and -r options" This reverts commit d94e034d504682be56fc2e9d20ac2c0fe15b70ec at the request of des@, as it seems to have broken the pass-through case.
rk_gpio: defer level-IRQ EOI until source line is driven low
The previous PIC bring-up (ccda002ca10) added pic_disable_intr,
pic_enable_intr, pic_pre_ithread, and pic_post_ithread, but omitted
pic_post_filter. Per the PIC contract pic_post_filter is non-optional;
a follow-up enforcement pass is planned that will panic() if any of the
three (pic_pre_ithread, pic_post_ithread, pic_post_filter) is missing.
This patch also fixes the EOI ordering for level-triggered IRQs (raised
by mhorne in the v1 review). Writing PORTA_EOI before intr_isrc_dispatch
is correct for edge pins, but wrong for level pins: the source device
has not yet deasserted the line, so the latch immediately re-arms and
the controller storms.
- rk_gpio_intr: EOI edge pins per-pin before dispatch (matches the
pre-patch behavior for the common case); for level pins defer EOI
to the post-dispatch path. Stray (no consumer) level pins still
get EOI'd here because no consumer will run to clear the source.
- rk_pic_post_filter: new method, EOI level pins after the filter
has read+cleared the source device's IRQ register.
- rk_pic_post_ithread: EOI level pins after the ithread has driven
the source low, before unmasking, so the chip latch is clean when
we re-enable delivery.
Shape mirrors tegra_gpio(4) (sys/arm/nvidia/tegra_gpio.c). No new
sysctls, no scaffolding.
Smoke-tested on RockPro64 (RK3399) with fusb302 INT_N (level-low GPIO
IRQ): IRQ rate steady at ~28/s under USB-C activity vs the 210 kHz
storm the original missing-mask bug produced.
Signed-off-by: Kyle Crenshaw <B1nc0d3x@gmail.com>
Reviewed by: mhorne
Fixes: https://cgit.freebsd.org/src/commit/?id=ccda002ca10f ("rk_gpio: implement PIC masking methods and mask unhandled IRQs")
Pull Request: https://github.com/freebsd/freebsd-src/pull/2245
Revert "rk_gpio: defer level-IRQ EOI until source line is driven low" There is a more correct / preferable scheme for handling of EOI. Requested-by: mmel This reverts commit 8ffb400bfd64102ac2a49639ccbbfffbe0c6f127.
tests/ktls: merge two sysctl checking helpers into one No functional change.
Revert "tests/ktls: merge two sysctl checking helpers into one" With certain sysctl configuration the test will fail. This reverts commit 801c0f383c0a719165c21ff5c29f231fb7b920c4.
This was a good idea, but we don't build metapackages in the kmods repo so it ends up breaking the release build. I might resurrect this change if/when the kmods repo includes the wifi-firmware-kmod metapackage. This reverts commit bda8028146694ee490543b35e3349e060936fde4. MFC after: 1 second
A native route Netlink interface will replace this stack. Requested by: glebius This reverts commit 1ccf543b21eff6e0828142e5c1d09519247143f4. This reverts commit 2c04cfa148ec4dd5cef7e228aaea6a05957fcb15. This reverts commit 2d6114f6d26bf7dfa5ad94e1db9b09ee7108dc7a. This reverts commit d15f2551b25f79ddcbe289faa95e655100b952da. This reverts commit ceb282bbd62eed5e84df9abaede0dd183f66997a. This reverts commit c30021fe0df9e045a17292dbe50dfc054b69871f. This reverts commit fb1820d23a04856a6d3047b4c088cc8df8f76da1. This reverts commit 8696cc600f44767e7988a92c8e6fb943e97d4cc7.
This reverts commit 45645518ea19ccb4761aee3a525aab2f323d37d4. Although this changed looks like it should just be a harmless change to bookkeeping, it turns out that it changes the termination condition of the initial device scan, resulting in it never finishing. This causes the boot to hang forever coming up. Since I don't have good access to hardware, I'm reverting until the exact details can be sorted out. Reported by: Edward Scroop Sponsored by: Netflix MFC After: 1 week
vtnet: move offload functions to virtio_net.h to share them Move the functions vtnet_rxq_csum() and vtnet_txq_offload() and the subfunctions they call from if_vtnet.c to virtio_net.h. This allows us to call these functions from if_tuntap.c and if_ptnet.c. virtio_net.h already contained a copy of these functions, but a copy of an outdated version. The functions evolved in if_vtnet.c. In if_vtnet.c, the copy has never been used because it increments counters in their own functions. This patch removes the outdated copy from virtio_net.h and moves the new version of the functions from if_vtnet.c to virtio_net.h. if_tuntap.c, if_ptnet.c, and if_vtnet.c just call these functions, and if_vtnet.c increments its counters depending on the return value. Reviewed by: tuexen MFC after: 1 month MFC to: stable/15 Differential Revision: https://reviews.freebsd.org/D57299
Revert "vtnet: move offload functions to virtio_net.h to share them" This reverts commit 44cddaa99dee0a634cf2713f71e799eb41397355. It breaks the LINT-NOIP config.
whereis(1): Respect PORTSDIR variable PORTSDIR is a very common variable that points to the location of a ports directory. Make whereis(1) respect this variable too. Reviewed by: arrowd@, christos@ Approved by: christos@ Differential Revision: https://reviews.freebsd.org/D56845
Revert "whereis(1): Respect PORTSDIR variable" This reverts commit edadc3f9051595a9c2e693d8ab666a50b9e7a21a.
contrib/lutok: remove autotools and doxygen generated files These files provide no value in the FreeBSD tree proper and change frequently, depending on what machine I generate the release tarball on (and what versions of autotools are on the host). Nuke the autogenerated files to avoid bloating commit history/the tree. MFC after: 3 days Requested by: Benjamin Jacobs <freebsd@dev.thsi.be>
Revert "contrib/lutok: remove autotools and doxygen generated files" I messed up the commit title. Redo it with a correct summary. This reverts commit 3dbe1800848ca71f5cc083f1d5c5b4759b601fae.
This reverts commit fc0a5ae094308609927e9c355928d08923278325. This changed the sysctl node path and broke IB user space tools (ibv_*). PR: https://bugs.freebsd.org/bugzilla/show_bug.cgi?id=298485 Reported by: Ariel Ehrenberg (aehrenberg nvidia.com)
virtual_oss(8): Fix cuse.ko check virtual_oss(8) checks if cuse(3) is loaded. However, kldload(2) ends up calling kern_kldload that checks permissions first. It is only later on in linker_load_module that -EEXIST is returned if the module is already loaded. That means that users that can't load modules, always get a -EPERM error first even if cuse.ko is already loaded and ready to use. Change it to check if the kernel module is already loaded and try load it if it isn't. In addition move the program's arguments parsing early on because otherwise, a user can't even access the program's help if cuse.ko is not loaded and the user doesn't have permissions to do it. Approved by: obiwac@ Differential Revision: https://reviews.freebsd.org/D59621
Revert "virtual_oss(8): Fix cuse.ko check" parse_options() was moved above cuse_init(), which makes every single regular virtual_oss invocation which uses cuse_dev_create() fail. This reverts commit f014795ec3bd5efb88dfc249599e9665dc10a59e.
This reverts commit e903655c50fe2e710a4cdfe6e638454012ef6324. The commit comes from the OpenZFS merge. OpenZFS introduced in d2f5cb3a5 two new libraries: libbtree and librange_tree. We don't have these in our build. To keep diffs as minimal as possible and avoid future breakages we should import these for the FreeBSD build with bells and whistles. After this issue is resolved, e903655c5 can be re-committed.
Not classified automatically, and waiting for manual attention.
-- no commits in this category this week --
Dates:
cgit.freebsd.org/src. Git accurately records the
order of commits, but not their dates.Automatic grouping:
This reverts commit \\b([0-9a-fA-F]{40})\\b
and the hash was found in this week's commits.
Automatic categories:
Source code:
Generated with commits-periodical 0.21 at 2026-10-05 16:55:06+00:00.
This work is supported by Tarsnap Backup Inc.
Alternate version: 2026-07-01 (debug) (contains info about the classification)