Today (2012-06-28) we are rolling out an unusually large security update. The update is necessary to fix known vulnerabilities. Unfortunately, the change has a small chance of breaking customer-compiled C extensions. We did our best to retain binary compatibility but we cannot guarantee it in every single case.
We ask all developers that have installed C extensions to verify that their programs are still starting up correctly. If not, a simple recompile should be sufficient to fix the problem.
Maintenance of switching infrastructure on 2012-07-05
On 2012-07-05 between 22:00pm and 23:00pm CEST we will change the power connection of one of our switches. This will cause a short interruption of network connectivity for some VMs. We expect the interruption to take less than 5 minutes.
This change is necessary to modernize the power distribution system within our racks. We try to perform this using redundant power connections without interruption of service as far as possible. During the next weeks we will also update our second rack which will require another short switch interruption. We will announce the required maintenance window separately.
Inconsistent PostgreSQL backups
Yesterday we found and fixed a critical bug in our backup server configuration.
A typo in the exclude file caused PostgreSQL write-ahead logs (WAL) not to be backed up. The PostgreSQL heap, which contains regular table data, has been backed up constantly though. This means that we were able to restore PostgreSQL data, but could not guarantee a consistent restore of transactions started between the last checkpoint and the time the snapshot for the backup was made.
The bug has been fixed yesterday. We have verified that all PostgreSQL servers' write-ahead logs have been backed up tonight so restores from backups performed since 2012-05-16 will not suffer from this problem.
A typo in the exclude file caused PostgreSQL write-ahead logs (WAL) not to be backed up. The PostgreSQL heap, which contains regular table data, has been backed up constantly though. This means that we were able to restore PostgreSQL data, but could not guarantee a consistent restore of transactions started between the last checkpoint and the time the snapshot for the backup was made.
The bug has been fixed yesterday. We have verified that all PostgreSQL servers' write-ahead logs have been backed up tonight so restores from backups performed since 2012-05-16 will not suffer from this problem.
Network maintenance on 2012-04-26 at 21:30-22:30 CEST
We will update our switch firmware on Thursday 2012-04-26 between 21:30 and 22:30 CEST (19:30-20:30 UTC). During these updates, we expect several short losses of network connectivity (<5min). Please note that customer VMs may be unreachable for short amounts of time.
The switch firmware updates will help us to ensure future network reliability and security. We apologize for any inconvenience.
The switch firmware updates will help us to ensure future network reliability and security. We apologize for any inconvenience.
Connectivity Issues 2012-04-10 (12:07pm-12:33pm CEST)
From 12:07pm to 12:33pm CEST we experienced a connectivity loss on our upstream connection to the data center. All physical and virtual customer machines have been affected.
The connectivity loss was caused by faulty switch hardware of our data center operator. The switch has been replaced immediately and connectivity has been restored.
We apologize for the service interruption.
Saturday, 2012-03-03 22:00 CET: Restart of all VMs due to security update
We need to perform a security update on our VM hosts which requires a restart of all virtual machines.
The restart will be performed beginning from Saturday, 2012-03-03 22:00 CET and will require about 4 hours until all VMs will have been rebooted. Reboot time of individual VMs may vary a lot depending on the need of performing a file system check. We expect most VMs to reboot within 10-20 minutes.
For details about the security issue, see http://www.gentoo.org/security/en/glsa/glsa-201202-09.xml).
Please excuse the short notice, according to our security policy we strive to install relevant updates as quickly as possible to minimize impact of known issues.
The restart will be performed beginning from Saturday, 2012-03-03 22:00 CET and will require about 4 hours until all VMs will have been rebooted. Reboot time of individual VMs may vary a lot depending on the need of performing a file system check. We expect most VMs to reboot within 10-20 minutes.
For details about the security issue, see http://www.gentoo.org/security/en/glsa/glsa-201202-09.xml).
Please excuse the short notice, according to our security policy we strive to install relevant updates as quickly as possible to minimize impact of known issues.
Connectivity issues 2012-01-29 (4:43pm-5:25pm CET)
From 4:43pm to 5:25pm CET we experienced high package loss on our upstream connection to the data center.
The cause of this was a malware infestation on a customer VM. No infrastructure components appear to have been compromised.
Connectivity has been reliable again since stopping the malware process.
The affected VM has been stopped and we are in contact with the customer to resolve the issue.
The cause of this was a malware infestation on a customer VM. No infrastructure components appear to have been compromised.
Connectivity has been reliable again since stopping the malware process.
The affected VM has been stopped and we are in contact with the customer to resolve the issue.
Subscribe to:
Posts (Atom)