Thursday, August 20, 2026
Incident Report #039: Backup Server Upgrade Fail
What are Incident Response Reports?
As a community‑operated and governed virtual internet exchange, FurrIX maintains
a commitment to open and honest communication with its members. During the normal
operations of the exchange, our network and its supporting systems may encounter
operational defects, bugs, failed changes or attacks on our infrastructure. When this
happens, the FurrIX volunteers publish an incident response report to ensure all members
and peers remain informed as to what happened, how it went down and what we did to
recover or resolve the issue. As a hobbyist‑rooted vIX, we aim to keep communication clear,
accessible and practical to the best of our ability.
What Happened
Our backup server had fallen significantly out of date across its software stack and during
an attempted upgrade. The update process broke roughly halfway through, leaving the system
in an unstable state and temporarily removing our ability to perform backups for the night.
Despite the failed upgrade, our volunteers successfully preserved all datastores before any data
loss could occur.
Upon assessing the severity of the failed upgrade, it became clear that attempting to salvage the
partially‑updated system would be too risky to FurrIX’s data. Our volunteers therefore used their
emergency governance override to authorize a full rebuild of the backup server. This decision allowed
us to move forward with a clean installation of the latest Proxmox Backup Server (PBS) version rather
than attempting to repair a heavily broken environment. Once the new PBS installation was in place,
the volunteers imported the preserved datastores and verified that the system could read all existing
backup data without any errors.
Points of Note
- Backups were unavailable for one night during the failure and rebuild window.
- No CT or VM backups appear to have been lost.
- All on‑hand data has been verified as readable on both PHYONE and PHYTWO.
- Normal backup operations have resumed under the rebuilt PBS environment.
The exchange avoided data loss and now has a fully updated and stable backup server. As always,
we appreciate the patience and understanding of our members and peers while we maintain and
improve the infrastructure that keeps FurrIX running.
Wednesday, August 19, 2026
Transparency Report #013 — Monitoring Anomalies and RRL Activity
What are Transparency Reports?
As a community‑operated and governed virtual internet exchange, FurrIX maintains
a commitment to open and honest communication with its members. From time to
time, operational work may occur that affects the exchange or its supporting infrastructure.
When this happens, the FurrIX operations team publishes a transparency report to
ensure all members remain informed. As a hobbyist‑rooted vIX, we aim to keep
communication clear, accessible and practical to the best of our ability.
What Happened
On the afternoon of 18 August 2026, our monitoring system (LibreNMS) began displaying
abnormally large bandwidth spikes for NS2, including inbound and outbound traffic peaks
exceeding hundreds of megabits per second. These values were inconsistent with NS2’s
QoS limited capacity and did not match host‑level counters or resolver behavior.
During the same period, NS2 received a high‑volume PTR sweep from the IPv6 prefix
2a05:f480:2400::/48. Bind’s Response Rate Limiting (RRL) immediately escalated and
began dropping all responses to the prefix. This prevented any outbound meaningful
traffic spikes on the VM and kept resolver load stable.
A review of physical host metrics, VM‑level counters and RRL logs confirmed that
NS2 never transmitted the large volumes of traffic shown in LibreNMS or displayed on
our website graphs. The anomaly was isolated to the monitoring layer and one of our
volunteers pointed out to us that our LibreNMS VM is hitching, slow to respond and
sometimes timing out.
Changes to the Exchange
Our volunteers identified two issues within the monitoring stack:
- Interface Mapping Drift:
LibreNMS occasionally identifies the wrong interfaces for some of our VMs.
We’re currently investigating what is causing this and have an idea, but do
not want to point fingers until we are sure.
- Poller Delays:
The LibreNMS VM showed signs of slow polling and missed intervals. The
internal logs are showing both missed polls and ICMP issues within the
whole exchange, but our testing and probing shows otherwise.
We believe these issues caused the monitoring system to display incorrect bandwidth
values for NS2 and a few other VMs. No changes have been made to the resolver or network
at this time. Operations is preparing remediation options, including a rebuild of the NMS VM since
it still contains data from the project we split from, and will be presenting them to oversight
for consensus before proceeding.
Are Exchange Operations Affected?
No.
NS2 continued operating normally throughout the event. RRL successfully mitigated
the abusive PTR sweep, QoS remained effective and host‑level counters showed
stable, low traffic consistent with expected resolver behavior. The only affected
component was the monitoring system’s visibility layer. Member‑facing, peering and
internal services were not impacted.
Sunday, August 2, 2026
[Incident Report #038][vIX] NS Abuse Impacting Exchange Stability
What are Incident Response Reports?
As a community‑operated and governed virtual internet exchange, FurrIX maintains
a commitment to open and honest communication with its members. During the normal
operations of the exchange, our network and its supporting systems may encounter
operational defects, bugs, failed changes or attacks on our infrastructure. When this
happens, the FurrIX volunteers publish an incident response report to ensure all members
and peers remain informed as to what happened, how it went down and what we did to
recover or resolve the issue. As a hobbyist‑rooted vIX, we aim to keep communication clear,
accessible and practical to the best of our ability.
What Happened?
During August 1st and August 2nd, 2026, the FurrIX vIX experienced a significant
service disruption caused by an unexpected surge of traffic hammering our name
servers. This traffic overwhelmed our exchange fabric, resulting in instability and
degraded performance for multiple services- up to the point of making the exchange
itself unreachable and completely saturated our link to our upstream.
During the events, we have been observing:
- vIX reachability issues — The virtual exchange entered a degraded state
due to excessive inbound NS load.
- NS1/NS2 saturation — Both name servers experienced heavy recursive query
pressure, exhausting logging capacity and reducing responsiveness.
- MA BGP instability — Member BGP sessions dropped from the vIX due to excessive traffic.
- Monitoring failures — NMS monitoring temporarily dropped during peak saturation.
- General service degradation — Web, email and auxiliary services were intermittently unreachable.
This event originated from external recursive DNS traffic and did not involve any FurrIX members.
Upon reviewing logs, traffic alerts and our periodic graphs, it was determined this may be automated
scanner traffic that found and exploited our infrastructure. Our volunteers are currently making
updated procedures and are in the process of updating the vIX configuration to deal with this.
What are we doing to fix this?
During investigation, it was determined that our existing rate‑limit configuration inside Proxmox
did not behave as we originally thought they would and thus did not protect the exchange fabric
from inbound saturation. Furthermore, the traffic that slapped the exchange was rotating through
IPv6 addresses at a rate and pattern that mostly avoided our BIND9 rate limits, which has been
eye opening in itself.
To address this, FurrIX is implementing the following corrective actions:
- Migrating all rate‑limiting from Proxmox to the OpnSense router VMs, where
shaping and policing can be applied correctly at the network demarcation.
- Reworking our NS policy, including tightening recursion rules, implementing
stricter query controls and improving protections for authoritative services.
- Enhancing network monitoring, adding visibility at the router, VM and hypervisor
layers to detect similar events earlier; including active alerting to our Discord.
Exchange operations may remain unstable while volunteers complete these changes and validate
the new configuration.
Current Status
As of this post, the exchange fabric has recovered and all core FurrIX vIX services are reachable.
Additional hardening work is ongoing and further updates will be posted once our volunteers have
completed all the ongoing work.
Wednesday, July 22, 2026
[Incident Report #037][vIX] Platform Issues and Tooling….
What are Incident Response Reports?
As a community‑operated and governed virtual internet exchange, FurrIX maintains
a commitment to open and honest communication with its members. During the normal
operations of the exchange, our network and its supporting systems may encounter
operational defects, bugs, failed changes or attacks on our infrastructure. When this
happens, the FurrIX volunteers publish an incident response report to ensure all members
and peers remain informed as to what happened, how it went down and what we did to
recover or resolve the issue. As a hobbyist‑rooted vIX, we aim to keep communication clear,
accessible and practical to the best of our ability.
Summary
On July 22nd, 2026 at approximately 12:45 EST, the FurrIX vIX entered a degraded
operational state. We lost access to PHYONE’s Proxmox control plane and experienced
partial failure of NS1’s DoT/DoH services. The root cause was a failure in newly
developed SSL certificate distribution tooling, which corrupted certificate stores on
both PHYONE and NS1. This resulted in service outages and exposed gaps in our
recovery lifelines.
What Happened?
Our lead engineer was developing new automation to handle SSL certificate
distribution across the vIX, aiming to reduce manual volunteer workload.
During a trial run on the distro network, the tooling behaved unexpectedly
and corrupted the certificate stores on both PHYONE and NS1.
Furthermore, when it was noticed that we no longer had UI access, our volunteers
tried using the recovery options we had thought were fully in place and found out
very quickly that neither lifeline was fully operational. For roughly two hours, the
vIX was running headless with limited administrative control.
We were seeing the following issues:
- Control Plane Damage — Proxmox UI on PHYONE offline
- NS Partial Failure — NS1’s DoT/DoH endpoints failing
- RNDC and other control‑plane operations to become unstable
The code appeared correct during review, but a permissions issue slipped through and
only manifested once deployed. A right fucken oops, that was. Compounding the issue,
when volunteers attempted to use our recovery lifelines, we discovered that neither of
them were fully operational. For roughly two hours, the vIX was running headless with
limited administrative control.
What did we do to fix this?
We learned that while our OOB recovery network was operational, our recovery shell
accounts had not been setup on PHYONE. We promptly reached out to the data center
for a KVM to be put on the machine as soon as possible. Upon getting that setup, our
volunteers deployed our recovery account via the machine shell and then detached from
KVM to continue the recovery as to our DRP for ‘Access Loss, PHYONE’ and we modified
our SSL handling scripts to ensure that NS1 set the correct permissions on the cert store
and that PHYONE now validates certs before installing them and reloading services.
During the recovery, we also learned the same issue was affecting tooling on NS1 itself and
we got that sorted to make sure it also validates its certs and keys before reloading any
services.
While the recovery was ongoing, we had a few blips in networking as the KVM was connected
and disconnected from PHYONE- but we should be fully alive and working again. Oh yea, we also
added Discord alerts to our scripts so that we are able to keep an eye on their execution and
catch any problems that might occur.
Response & Recovery Actions
1. Restoring Access to PHYONE
- Verified OOB recovery network was operational
- Discovered recovery shell accounts were not configured
- Contacted the data center and requested KVM attachment
- Once KVM was online, deployed recovery account via local shell
- Detached KVM to minimize network blips and continued
recovery per DRP “Access Loss, PHYONE”
2. Repairing Certificate Store PHYONE
- Identified corrupted cert stores on PHYONE
- Updated SSL handling scripts to:
— Validate certs before installation
— Validate key/cert pairing
— Reload services only after validation passes
3. Fixing NS1 Tooling and Cert Store
- Found identical validation/permission issues affecting NS1’s tooling
- Updated scripts to ensure NS1 validates certs and keys before reloading services
- Apply correct permissions (bind:bind)
Root Cause
A permissions and validation oversight in new SSL distribution tooling caused certificate
corruption on PHYONE and NS1. Lack of fully configured recovery accounts delayed restoration.
Lessons Learned
Going forward with the development and upkeep of the FurrIX vIX,
our volunteers will be applying the following lessons to future scripts
and any custom tooling:
- Automation touching cert stores must validate before overwrite
- Permissions must be explicitly set every time
- Recovery accounts must be deployed and tested on all hypervisors
- Monitoring hooks (Discord alerts) are essential for early detection
For onlookers wondering why this was not sandboxed more aggressively before
being put into production, it is a hard truth that FurrIX does not have a replicated
environment offsite for testing these kinds of tools. Meaning a lot of the time, we
are heavily crawling our own tooling before we deploy and sometimes not so obvious
issues can crop up that our volunteers haven’t thought about before hand.
Wednesday, July 15, 2026
[Incident Report #036][DC] Network Cabling Issues….
What are Incident Response Reports?
As a community‑operated and governed virtual internet exchange, FurrIX maintains
a commitment to open and honest communication with its members. During the normal
operations of the exchange, our network and its supporting systems may encounter
operational defects, bugs, failed changes or attacks on our infrastructure. When this
happens, the FurrIX volunteers publish an incident response report to ensure all members
and peers remain informed as to what happened, how it went down and what we did to
recover or resolve the issue. As a hobbyist‑rooted vIX, we aim to keep communication clear,
accessible and practical to the best of our ability.
What Happened?
On July 14th, 2026 at around 2045EST, the FurrIX vIX went offline for about thirty-five
minutes. The volunteers at FurrIX reached out to the data center for some insight after
doing our own troubleshooting. It was discovered that our secondary server was still
alive but our primary had no response.
We were seeing the following issues:
- vIX reachability — The virtual exchange was temporarily in a degraded state.
- NS1 Failure — Members were relying on NS2’s zone cache temporarily.
- MA BGP Failures — Our BGP sessions with members temporarily failed.
- Web/email Failures - These services stopped responding.
- Games-3P Failure - Game services went offline.
What did we do to fix this?
We contacted the data center to determine the scope of the event and were informed
that the techs at the upstream data center had found a loose network cable and that
was the cause of our server going offline. They have reseated the cable and the exchange
is now back online and services are reachable once again.
As of this post, all FurrIX vIX services have recovered.