eScience| UCloud: DeiC Interactive HPC Operational |
| — Provider: Bitten Operational |
| — Provider: Odense Operational |
| Hippo: Local Slurm Cluster Operational |
We are rolling out important security patches on the system, which requires all machines to be rebooted. Running jobs will be terminated.
Jul 10, 13:17 - Jul 12, 14:53The remaining compute nodes were rebooted during the afternoon.
Jul 12, 14:49We are updating the SDU eScience Center servicedesk so it is currently unavailable.
Jun 30, 09:00 - Jun 30, 09:50The Bitten system is currently unavailable. During the night a transformer station in Aabenraa exploded and caused widespread problems for the surrounding power grid, including our data center. There has been no loss of data, but we are keeping the system offline during the weekend to asses the status and potential damages.
Jun 26, 23:30 - Jun 29, 07:23Storage is now accessible, but compute nodes are still offline.
Jun 28, 10:10Compute nodes are now back in production and everything should be in working order.
Jun 29, 07:23Around 5:30 this morning all machines at the Bitten provider were unexpectedly rebooted. We are working on bringing systems back online.
Jun 18, 05:30 - Jun 18, 07:57We are currently experiencing issues with the SDU eScience Servicedesk. Tickets submitted via the UCloud support form and the SDU eScience support email are not reaching the Servicedesk. Longer response times are to be expected for now. It is still possible to open and reply to tickets, but only directly via support.escience.sdu.dk.
Jun 3, 06:00 - Jun 4, 10:13We have scheduled an extraordinary maintenance window with short notice, because we need to rollout some important changes to the system. All compute nodes will be rebooted during the maintenance window and they might be unavailable for up to two hours. All running jobs will be terminated.
May 23, 07:00 - May 23, 08:29Maintenance has been completed and all nodes are now available again.
May 23, 08:30Our ISP has announced a maintenance window on Friday (May 8th) that very likely will disrupt our internet connection to Bitten.
May 8, 06:00 - May 8, 13:38SDU is currently experiencing a power loss, which is also affecting services in our data center. This is also affecting the storage system at Bitten, which cannot talk to the second tier at the SDU site.
May 6, 16:10 - May 6, 21:00We are experiencing some stability issues with the new cpu-amd-zen5 machines. As a result, they sometimes need to be rebooted, affecting all jobs running on the node. We are working on finding a solution.
May 6, 15:32 - Jun 3, 08:04We temporarily lost our internet connection to the new data center. The problem has been identified by the ISP and we are waiting for them to solve it.
May 5, 15:23 - May 6, 05:05The connection was finally restored again around 5am.
May 6, 05:20We are moving the UCloud platform to a new data center and this will require extended downtime. Downtime will start on Monday, April 27th, and continue for up to one week. The system is expected to be back online Monday, May 4th, at the latest.
Apr 27, 08:00 - May 4, 07:00We are now taking UCloud offline for the data center migration.
Apr 27, 07:59The migration has been completed and the system is back online.
May 4, 07:01Unfortunately we have had to terminate background tasks (file transfers and copy operations). This was needed to deploy an updated which should improve the stability of the same feature. The background tasks have to be resubmitted to continue. We apologize for the inconvenience.
Mar 17, 10:40 - Mar 17, 10:40UCloud was down from 01:00 until 08:55 due to a bug in the UCloud code. We have temporarily disabled the broken code and are working on deploying a permanent fix. UCloud is expected to work normally while we work on a fix.
Feb 1, 01:00 - Feb 1, 08:55After an update of the Nvidia drivers we are now experiencing problems with the Nvidia H100 (u3-gpu) machines. We will start downgrading the drivers again next week to restore full functionality.
Jan 31, 11:32 - Feb 4, 07:47Drivers have been downgraded on all machines that experienced problems.
Feb 4, 07:48We are aware of an issue causing public IPs to not correctly be attached to machines. We are working on a fix.
Jan 29, 07:04 - Jan 29, 09:30The issue has been resolved.
Jan 29, 09:30UPDATE: This maintenance window has been expanded and it now covers ALL services offered by SDU eScience.
We will be performing maintenance on January the 28th between 12:00 and 20:00. UCloud will be down during this period and jobs on SDU/K8s and AAU/K8s will be terminated at the start of the maintenance window.
Jan 28, 12:00 - Jan 28, 17:27First part of the maintenance has been completed. We are now working on deploying the new version of UCloud.
Jan 28, 15:20Update has completed. Please report errors that you may find.
Jan 28, 17:28On December 1st there will be scheduled maintenance in the SDU data center, which requires a complete shutdown of all servers. For this reason UCloud and all other services offered by SDU eScience will be unavailable during the entire working today. We expect systems to be back online late in the afternoon.
Dec 1, 07:00 - Dec 1, 15:16All services should be back up and running.
Dec 1, 15:16The u2-gpu machines are currently experiencing hardware problems. User jobs are able to run, but they can be killed at any point due to maintenance.
Nov 18, 09:00 - Dec 3, 11:32The machine has been powered off and hardware replacements should be performed later today.
Dec 3, 09:21The hardware has been replaced and all GPUs are working again.
Dec 3, 11:33UCloud is currently experiencing issues, we are working on fixing the problem.
Nov 17, 19:26 - Nov 17, 20:21During the night UCloud had an internal issue, which made it impossible to access jobs and files.
Nov 14, 00:12 - Nov 14, 06:48UCloud has been updated with the latest round of bug fixes and improvements to the UI. As always, this may have caused a few minutes of disruption to the service. Sorry for the inconvenience.
Nov 11, 09:37 - Nov 11, 09:36On Sunday, October 26th, several UCloud jobs that have been running for more than 14 days will be terminated due to hardware maintenance.
Oct 26, 10:00 - Oct 26, 18:07The jobs have been terminated.
Oct 26, 18:07The SDU/K8s provider was unavailable between 12/10/25 15:56:34 and 12/10/25 17:04:25 due to a software bug. A bug fix has been released now (13/10/25 07:00).
Oct 12, 15:56 - Oct 12, 17:04UCloud is being restarted for an update. The update will take a few minutes.
Sep 30, 08:50 - Sep 30, 08:52Two of the H100 nodes were rebooted this morning due to hardware maintenance, the first one around 9.00 and the second one around 10.30 A couple of jobs where stopped during the reboot.
Sep 26, 08:00 - Sep 26, 11:05Around 8:50 this morning we started decommissioning an old storage system, which unfortunately affects the SDU/K8s compute nodes, making them partially unresponsive.
Sep 22, 08:52 - Sep 22, 10:45Things should finally be returning to normal. If jobs are stuck for more than 10 minutes, start a new one.
Sep 22, 10:31A fluctuation in the power grid caused around 30 machines to reboot in the UCloud server room. Jobs running on the machines were terminated during the reboot.
Sep 9, 21:45 - Sep 10, 06:30The machine 'nodea0-19' has been rebooted due to an error with one of the GPU cards. We are monitoring the node for a potential hardware issue.
Sep 2, 07:35 - Sep 19, 13:19The error has reappeared, the card will most likely need to be replaced.
Sep 4, 07:48A support case has been opened to get the card replaced.
Sep 10, 07:49Hardware maintenance has been initiated.
Sep 19, 12:46The faulty GPU has been replaced.
Sep 19, 13:19Between 00:21 and 08:00 UCloud was running with elevated error rates causing the job page to not load correctly. The issue has been resolved.
Aug 29, 00:21 - Aug 29, 08:01We are aware of a situation causing jobs to not start. We believe this is related to the storage system. We are currently investigating the situation.
Aug 20, 12:26 - Aug 20, 15:22Projects receiving storage allocations from "Type 1 - KU" for SDU/K8s were temporarily unavailable. The allocations should be available once again now.
Aug 15, 11:03 - Aug 15, 11:02We are aware of an issue causing usage to be tracked incorrectly for storage. We will restart the system to fix this issue.
Aug 14, 11:27 - Aug 14, 11:39We have completed a reset of usage numbers in storage which we believe will fix the issue. The system is now back online after a few minutes of downtime.
Aug 14, 11:39We are experiencing slower storage performance on some nodes, which are leading to jobs being slow to start. Jobs that would ordinarily start within a minute can now take up to a few minutes before they start. We are monitoring the situation.
Aug 14, 08:54 - Aug 15, 13:07An issue has been resolved leading to a crash is under investigation
Aug 13, 17:30 - Aug 13, 17:55Syncthing has temporarily been disabled while we investigate an issue with the system. We do not believe that the system will be re-enabled today (13/08/25). We hope to have an update with ETA or fix tomorrow (14/08/25).
Aug 13, 15:58 - Aug 15, 13:07We are currently testing a fix internally and hope to deploy the fix tomorrow (15/08/25).
Aug 14, 14:08Syncthing has been re-enabled. We will monitor the situation. If instability is reintroduced then we may need to disable it again. Status updates will be posted here if this becomes needed.
Aug 15, 13:08We are investigating an issue related to jobs not starting.
Aug 13, 14:52 - Aug 13, 15:58The is no longer present. We will continue to monitor the system.
Aug 13, 15:58UPDATE: The maintenance has been moved from the 12th of August to the 13th of August.
The SDU/K8s (DeiC Interactive HPC, SDU) service provider will be down on the 13th of August (13/08/2025). All jobs will be killed prior to the maintenance and will not be automatically restarted. The maintenance is expected to take place between 08:00 and 16:00.
Aug 13, 08:00 - Aug 13, 13:45Preparation phase of the maintenance took longer than expected. The primary work has started now and the provider is no longer accessible as noted in the original maintenance notice.
Aug 13, 09:56Maintenance has been completed. We will be monitoring the system over the coming hours and days.
Aug 13, 13:45We are investigating an issue which causes the "Open interface" button to not appear on the SDU/K8s provider.
Aug 8, 08:50 - Aug 8, 09:03The issue has been resolved.
Aug 8, 09:03We are looking into issues with UCloud, which causes most funtionality to not be working.
Aug 7, 21:15 - Aug 8, 08:50The problem has been identified and the system is operational again.
Aug 8, 08:45The AAU/K8s (DeiC Interactive HPC, AAU) service provider will be down on the 5th of August (05/08/2025). All jobs will be killed prior to the maintenance and will not be automatically restarted. The maintenance is expected to take place between 08:00 and 16:00.
Aug 5, 08:00 - Aug 5, 10:29Maintenance has completed and the system is available. We are monitoring the system for possible bugs.
Aug 5, 10:29