Afrihost - Pure Fibre Feedback Thread Part 2

Unfortunately, the reboots did not resolve the issue. Frogfoot maintenance teams are en-route to the data centre with replacement hardware. No ETR is currently available.

Update. The replacement is faulty and techs are sourcing another replacement from their warehouse.

-- The replacement line card which was installed turned out to be faulty. We are in the process of sourcing a new one from the warehouse.
 
Update. The replacement is faulty and techs are sourcing another replacement from their warehouse.

-- The replacement line card which was installed turned out to be faulty. We are in the process of sourcing a new one from the warehouse.
Thanks for the update.
 
Update. The replacement is faulty and techs are sourcing another replacement from their warehouse.

-- The replacement line card which was installed turned out to be faulty. We are in the process of sourcing a new one from the warehouse.
That reminds me of the auto-electrician appie. He went to the boss with a box of blown fuses; "Hey boss, they sent you a box of dud fuses", he said, after testing them one for one in the old car with a full short!!
 
Last edited:
*Outage Update*

*Outage Title:* Outage | CPT Subset | Degraded
*Outage Ref(solid):* FRG7013707
*Impact Start Time:* 2024-12-31 23:10:00
*Impact:* Degraded

*Outage Update:*
*2025-01-01*
-- Logged to T3 team for investigation
-- Teams are currently investigating impact.
-- Core team has been alerted to investigate further. More updates to be shared
-- Busy renaming routing instance for additional flooding that the team is seeing on the network.
-- Tier 3 team still seeing flooding on the network, Core is assisting with investigating further.
-- We are observing an ongoing flap on the Juniper, our core team is currently investigating
-- Techs are onsite hard rebooting the junipers to try and resolve the issue
-- The reboot of R2 is now complete. R1 has just restored, we are waiting for R2 to fully stabilize before bringing the interfaces up. Thank you for your continued patience as we work through this process. We will keep you updated on further developments
-- The spare line card has been collected at our office and is enroute to site for installation
-- Team is onsite with the line card waiting on core team to assist with the installation, we are pending ETTR for the installation. Further updates will be shared
-- The tech is busy changing the line card
-- The replacement line card which was installed turned out to be faulty. We are in the process of sourcing a new one from the warehouse.
-- Techs have received the replacement line card and are en route to site
-- Tech is on site, installation will commence shortly
-- Line card has been installed. We are pending Tier 4 to complete configurations and to see if services are restored and stable
-- ISPs are monitoring services and sharing samples of the clients that are still impacted
-- T4 is working through the samples to isolate the issue

*2025-01-02*
-- The fault is still being investigated. There is currently a joint troubleshooting with the hardware vendor. We are awaiting an update on the outcome of the investigation.
-- Sourcing an update from the teams. Feedback will be shared as soon as possible.
-- Investigation still ongoing. Fault has been escalated internally on juniper side.
-- Teams are in the background continuing with isolation and resetting individual routing instances that are showing flooding with the aim of reducing service impact.
 
*Outage Update*

*Outage Title:* Outage | CPT Subset | Degraded
*Outage Ref(solid):* FRG7013707
*Impact Start Time:* 2024-12-31 23:10:00
*Impact:* Degraded

*Outage Update:*
*2025-01-01*
-- Logged to T3 team for investigation
-- Teams are currently investigating impact.
-- Core team has been alerted to investigate further. More updates to be shared
-- Busy renaming routing instance for additional flooding that the team is seeing on the network.
-- Tier 3 team still seeing flooding on the network, Core is assisting with investigating further.
-- We are observing an ongoing flap on the Juniper, our core team is currently investigating
-- Techs are onsite hard rebooting the junipers to try and resolve the issue
-- The reboot of R2 is now complete. R1 has just restored, we are waiting for R2 to fully stabilize before bringing the interfaces up. Thank you for your continued patience as we work through this process. We will keep you updated on further developments
-- The spare line card has been collected at our office and is enroute to site for installation
-- Team is onsite with the line card waiting on core team to assist with the installation, we are pending ETTR for the installation. Further updates will be shared
-- The tech is busy changing the line card
-- The replacement line card which was installed turned out to be faulty. We are in the process of sourcing a new one from the warehouse.
-- Techs have received the replacement line card and are en route to site
-- Tech is on site, installation will commence shortly
-- Line card has been installed. We are pending Tier 4 to complete configurations and to see if services are restored and stable
-- ISPs are monitoring services and sharing samples of the clients that are still impacted
-- T4 is working through the samples to isolate the issue

*2025-01-02*
-- The fault is still being investigated. There is currently a joint troubleshooting with the hardware vendor. We are awaiting an update on the outcome of the investigation.
-- Sourcing an update from the teams. Feedback will be shared as soon as possible.
-- Investigation still ongoing. Fault has been escalated internally on juniper side.
-- Teams are in the background continuing with isolation and resetting individual routing instances that are showing flooding with the aim of reducing service impact.
Thanks for the update, CD.
 
*Outage Update*

*Outage Title:* Outage | CPT Subset | Degraded
*Outage Ref(solid):* FRG7013707
*Impact Start Time:* 2024-12-31 23:10:00
*Impact:* Degraded

*Outage Update:*
*2025-01-01*
-- Logged to T3 team for investigation
-- Teams are currently investigating impact.
-- Core team has been alerted to investigate further. More updates to be shared
-- Busy renaming routing instance for additional flooding that the team is seeing on the network.
-- Tier 3 team still seeing flooding on the network, Core is assisting with investigating further.
-- We are observing an ongoing flap on the Juniper, our core team is currently investigating
-- Techs are onsite hard rebooting the junipers to try and resolve the issue
-- The reboot of R2 is now complete. R1 has just restored, we are waiting for R2 to fully stabilize before bringing the interfaces up. Thank you for your continued patience as we work through this process. We will keep you updated on further developments
-- The spare line card has been collected at our office and is enroute to site for installation
-- Team is onsite with the line card waiting on core team to assist with the installation, we are pending ETTR for the installation. Further updates will be shared
-- The tech is busy changing the line card
-- The replacement line card which was installed turned out to be faulty. We are in the process of sourcing a new one from the warehouse.
-- Techs have received the replacement line card and are en route to site
-- Tech is on site, installation will commence shortly
-- Line card has been installed. We are pending Tier 4 to complete configurations and to see if services are restored and stable
-- ISPs are monitoring services and sharing samples of the clients that are still impacted
-- T4 is working through the samples to isolate the issue

*2025-01-02*
-- The fault is still being investigated. There is currently a joint troubleshooting with the hardware vendor. We are awaiting an update on the outcome of the investigation.
-- Sourcing an update from the teams. Feedback will be shared as soon as possible.
-- Investigation still ongoing. Fault has been escalated internally on juniper side.We have identified a bug that is causing the impact, we are going to upgrade the 2 Junipers in cpt.
We are currently planning for the upgrade. updates to follow.
-- Teams are in the background continuing with isolation and resetting individual routing instances that are showing flooding with the aim of reducing service impact.

Frogfoot has identified a bug that is causing the impact, they are going to upgrade the 2 Junipers in cpt.
Currently planning for the upgrade. updates to follow.
 
Frogfoot has identified a bug that is causing the impact, they are going to upgrade the 2 Junipers in cpt.
Currently planning for the upgrade. updates to follow.
Hi cavedog. Thanks for the update.... Also the rant below is not directed at you but at Frogfoot.

A bug is software or firmware related. So did they do an upgrade on the 31st just after 11pm when it started? Because Network Engineers that were only dispatched the next day should have known about the change then. Also what I don't understand is, according to Frogfoot, they only rebooted the equipment yesterday afternoon, then decided it was a line card and no this morning they want to upgrade (perhaps revert the previous upgrade). Change control dictates a Regression Plan if the change is not successful, where is that?

If this wasn't a Change, but a hardware problem why are complete Unit Spares not available? This is affecting the whole Western and Eastern Coastline...It's a critical failure point and should be covered by spares.
 
Frogfoot has identified a bug that is causing the impact, they are going to upgrade the 2 Junipers in cpt.
Currently planning for the upgrade. updates to follow.
*Outage Update*

-- The upgrade developed jointly with the Vendor has been prepared and is currently undergoing final checks. Once ready we will begin the upload to the affected devices which will take around an hour per device (2 devices in total). We will provide more accurate timeframes as we begin the processes.
 
I guarantee this is as frustrating to them as it is for you.

There are obviously obstructions causing the delay in having your issue sorted out on the FNO side.

@AfriNatic as advised if he could he would by all means have you activated and running within seconds if he could.

I understand your frustration and how long it has taken to be resolved but if one side is not playing nice you cannot blame the other side that is trying their best to help.

@AfriNatic has many times gone out of his way to help, as far as even trying to help a certain FNO in my area to configure their equipment correctly whilst on leave and in my area.

I really hope that you can be up and running asap.
3 weeks gone. No change. No attempt at an actual explanation.
 
*Outage Update*

-- The upgrade developed jointly with the Vendor has been prepared and is currently undergoing final checks. Once ready we will begin the upload to the affected devices which will take around an hour per device (2 devices in total). We will provide more accurate timeframes as we begin the processes.
*Outage Update*


-- Team is still busy with an internal investigation
-- We are sourcing the latest updates from the team
-- Teams is still finalizing the checks for the upgrade developed with the Vendor. Some risks have been identified which could cause further impact and the Vendor has added a more senior engineering team to assist with this troubleshooting. More updates will be provided in the next 30 - 45 minutes.
-- Additional logs collected by the Vendor indicate that there is an underlying issue on the FPC's that needs to be resolved before any firmware upgrade is implemented. It is highly likely that resolving this underlying issue will also resolve the intermittency issues and negate the need for any additional firmware upgrades. Vendor and Frogfoot Engineering Teams are actively troubleshooting and we will provide more updates as they progress.
 
-- Additional logs collected by the Vendor indicate that there is an underlying issue on the FPC's that needs to be resolved before any firmware upgrade is implemented. It is highly likely that resolving this underlying issue will also resolve the intermittency issues and negate the need for any additional firmware upgrades. Vendor and Frogfoot Engineering Teams are actively troubleshooting and we will provide more updates as they progress.

-- T3 is pulling logs to continue to investigate. More updates to follow.

-- We are sourcing the latest updates from our T3
 
Packet loss on Frogfoot Fibre via Afrihost is a pulsing 100% every two seconds or so still in CPT area. Sad state of affairs. 😑
 
Packet loss on Frogfoot Fibre via Afrihost is a pulsing 100% every two seconds or so still in CPT area. Sad state of affairs. 😑


*2025-01-03*
-- New hardware devices have been installed and CORE teams are busy setting up and installing configs. Once completed we will perform a reboot under a more balanced load and then perform checks.
-- New device configurations are still underway, some services have been moved over to the new routers and we are expecting to perform a reboot once we have completed enough to balance the load. Further updates will be provided as we progress.
-- The Core team has completed moving over some services to the new routers and are now performing a reboot of the routers, once the reboot has been completed they will conduct some checks
-- The routers have been rebooted and the Core team is busy performing some checks. Further updates will be shared.
-- We discovered that there is an the underlying fault that triggered the juniper issues which had not yet been identified. Our change last night resolved the juniper issues, and with the junipers stabilized we have managed to identify the underlying issue. Core team is now working on resolving the underlying issue.
-- We have identified a loop on the network which seems to be originating from the table view switch. The core team have disabled that backhaul as a temp fix. We are monitoring for stability. A tech is en route to site and the ETA is 20minutes. Core team will work remotely with the tech to implement a permanent fix.
-- Tech has been on site since 9.30 am and the core team has been investigating remotely, we are awaiting further updates on the investigation.
-- The on-site tech and our CORE team have rebooted the switch, they are currently completing the required checks. Stability will be monitored over the next hour and we will provide further updates.
 
Packet loss on Frogfoot Fibre via Afrihost is a pulsing 100% every two seconds or so still in CPT area. Sad state of affairs. 😑
How is the packetloss now please test and advise
 
Top
Sign up to the MyBroadband newsletter
X