Showing posts with label IBN. Show all posts
Showing posts with label IBN. Show all posts

Thursday, January 5, 2023

An Architectural Perspective on Hierarchy In IP Networks - A Complex Puzzle Comprising Protocols, Topologies, Addressing & Systems (A Short Post)

 


If you ever bump into a Network Design book, course, blog or a webinar - most likely you are going to get introduced to this interesting & an important concept of " Network Hierarchy aka Hierarchical Networks Design Principle ".

Now depending upon which study materials and authors you follow, you would likely to come across different view points in terms of it's needs, pros & cons. Which in general not only contributes into more confusion among audience but also when I speak around with experiences Network Architects & Design Engineers, I often find them:

1. Having different interpretations of this concept and different view points

2. Considering this to be a very theoretical concept which you are likely to encounter in most Architecture & Design books but don't know about:

A. How to practice it (By applying theory to practice) 

B. How to measure it 

C. Missing the deep understanding of the topic at hand beside failing to understand its tradeoffs

So Idea behind this post is to offer you some architectural decision pointers to think through the problem statement and break it down into few tangible pieces by following another important network design principle " Separate the Complexity from the Complexity - Russ White " beside examining the rule 8 from RFC-1925

One of the model I personally always find handy is the SOS model from Russ White and you can use it too as a good ref. point.


So here is the quick list for you to think through in a more pragmatic manner:

1. What problems are you really trying to solve by introducing hierarchy into the Network (Go beyond theory) ?

2. Is it always possible to follow hierarchy? Specially in brown fields or during transitions (Think of old gear with still some lifetime left, mergers )

3. What are the downsides of introducing hierarchy ? (What harm it can cause and tradeoffs such us downgrade Agility, Flexibility, Organic Growth etc.)

4. Difference between Hierarchy vs. Symmetry vs. Modularity vs. Abstractions 

5. Different types of hierarchy/ layered approach to it ( physical level hierarchy,  logical level hierarchy,  hierarchy in addressing scheme, Protocol Level Hierarchy (ISIS Levels & Addressing ?) and so forth)

6. What data points you have in place to test your hypothesis to measure its impact on network

7. How these concepts are applied to different network environments - Enterprises (Campus <Wired and Wireless>, WAN/SDWAN, DC) vs Teclos vs CDNs vs Cloud Providers vs Web Scales vs Within public cloud virtual DC + Controller vs. Controller Less Architectures

8. Impact of introducing hierarchy on Visibility, Reporting and Performance mgmt. of the network

9. Impact on hierarchy on information hiding <reachability information> vs. topological information hiding (Aggregation vs. Summarization)

10. How all these choices will flow into your equipment sizing and potentially have an impact on your decision process

11. How will you apply all these concepts in a IPv10 network (IPv4 + IPv6 aka Dual Stack)

12. And if you are still brave enough :) , read through the further readings list to get to the bottom of this rat hole

Further Readings:

P-FatTree: A Multi-channel Datacenter Network Topology

Enabling Wide-spread Communications on Optical Fabric with MegaSwitch

Abstraction in Networks with Russ White

Hierarchical Network Design Overview

Engineer Versus Complexity

Optimal Routing Design

Navigating Network Complexity

Network Topologies

Five Number Summary for Network Topologies

Scaling MPLS Networks

The Side Effects Of Route Summarization

Avoid Summarization in Leaf-and-Spine Fabrics

Valley-Free Routing

Intra-Spine Links in Leaf-and-Spine Fabrics

Nonblocking versus Noncontending

Hierarchical IP Address Design and Summarization

Hierarchical IPv4 Framework

Fabric versus Network: What’s the Difference?

Liskov Substitution and Modularity in Network Design

Dragonfly+: Low Cost Topology for Scaling Datacenters

Reliability Basics- Part1

Network Centrality and Robustness

Swimlanes, Read-Write Transactions and Session State

Fifty Shades of High Availability

HTH...

A Network Artist ðŸŽ¨

Wednesday, January 4, 2023

Overlay Networks & Protocols Tradeoffs - Aka SDN aka IBN aka Magic aka Silver Bullet

 


A long time ago I wrote a short article on what really went wrong with SDN, now a few years later the topic still pops up in a conversation with the great Ivan & he acknowledged my list of Tradeoffs (things to watch out for carefully) related to overlay networks and protocols which seems to be de-facto standard for most modern Network solutions that we see around in Enterprises & Telcos.

  • Impact of overlay networks on visibility, reporting and performance management
  • Additional control plane that would result in additional abstraction layers and interaction surfaces and hence cascading effect in many situations
  • Impact on troubleshooting: how many solutions do we see in the market that can correlate underlay and overlay problems?
  • When it comes to sizing equipment in terms of control plane or data plane, it poses a new level of complexity an architect would need to deal with and in most cases vendors themselves won’t be able to offer much help in general rather than just asking you to believe in their words
  • I see lot of VXLAN and EVPN preachers, but let’s agree that mapping VLAN to VXLAN on 1:1 basis tells me you don’t know your stuff and believe too much in vendor marketing
  • EBGP underlay with IBGP overlay…man we can do better
  • Stitching two EVPN DCs with MPLS and SR: most of the implementations that I have seen were too complex and too fragile and thus results in a complex “policy.”

Further Readings:

Disjoint Path Routing and LP

HTH...

A Network Artist ðŸŽ¨

Wednesday, December 22, 2021

The QOS Fallacies & Failures in a Modern Hybrid IT World - Part 1 of 2

 


When I wrote about QOS the last time around 7 years ago, I must say I had high hopes from SDN & IBN as both of the paradigms were still evolving. Meanwhile there have been some unsuccessful attempt to automate QOS by throwing some sort of controllers into the mix by few vendors beside some others claiming they can solve this problem with the mighty IBN.

As we are about to move into 2022, many still wonder if QOS makes any sense at all in the context of modern networking ?

In order to find the answers, let's break the problem into two parts:

1. Why QOS has been so unsuccessful historically
2. What are our options moving forward

So let's focus on point 1 to begin with.

1. How do we get started ? - Interestingly enough over a dozen books have been written on QOS over the last 2 decades or so in the context of IP networking which mostly talks about details such as congestion management vs. congestion avoidance and so forth. But very few of them actually jumps into platform specifics in terms of capabilities and dependencies (both HW & SW). More interestingly I personally haven't come across a single QOS book myself yet which gives you any practical advice or framework/methodology around how to gather technical requirements in reference to Applications in order to plan and craft a QOS policy. So often I have seen people struggling to come up with one and given most QOS deployments are tactical rather than strategic, people often have time constraints to come up with a one in a short time. That's why many times people end up coping some references from recommended design guides etc. which hardly works in real life (Unless you were too lucky !).

2. Benchmarking & Capacity Mgmt. - Most small, medium & even couple of the large enterprises that I have worked with including Telcos don't seem to have both of these as mature practices in place. Benchmarking is though one of the key exercises you need get through to craft a good QOS policy beside being a necessity in any of a mature Capacity mgmt. framework/practice. The other problem you may likely to run here is that in order to do effective benchmarking & capacity mgmt. you need to invest into additional visibility & performance mgmt. tools to gather the required details which are usually quite expensive beside that fact that you need to train your team on tools and required operating skills (for example statistical analysis, Time Series, Sampling details etc.). Certain times you are likely to run into the problem where in the given tool may not be able to offer you reporting/data that you need natively which means you are always dependent on tool vendor about if their product road-map is aligned to your priorities and timelines. And be careful about if the tool allows you to run custom reports or exports required data in format you need. So if your capacity mgmt. is still on excel sheets, you know where you are heading.

3. Measuring latency incorrectly - This is perhaps more common than you might have thought beside that fact that most QOS books don't offer any practical advice here too and details around how latency needs to broken down across the spectrum. Can you tell me the breakup of end to end latency (Server - Client) and that too on hop by hop basis ?

4. QOS Lifecycle Mgmt. - While this area has improved a bit when it comes to modern networking gear, assuming majority of equipment still out there are old ones which doesn't offer much when it comes to QOS lifecycle mgmt. that includes Plan, Design & Implementation. But more importantly what they lack are capabilities such as QOS monitoring & reporting in real time beside the correlation with network health & events. After all you don't want to hire someone today to type couple of show commands in every few minutes and running the scripts won't be that helpful either for most part when you are dealing with scale. Again there are couple of commercial and open source tools available but its an exercise which takes time and resources beside that fact you should know exactly what you are looking for in which scenario. BTW...will that resource be from planning team, tools team or ops team is what I leave for you to figure out in real life. :)

Also with the rise of modern solutions which are mostly built around the magical controllers and overlays, you must think about QOS FCAPS capabilities in such environment carefully. For example while from overlay protocol perspective everything might be just a single hop away, the packet eventually still gets passed through the physical world in underlay. So its importantly to find out early if QOS policies will be:

- Static or Dynamic in nature (Given you are using controller of some sort)
- Correlation of QOS statistics between underlay & overlay
- How policy gets propagated 
- How controller interacts with other systems and policies for introducing dynamic behavior 
- Does the system allows Time Based QOS policies (interestingly enough most don't yet)
- Dummy policy dry run capabilities if supported 
- How your QOS policy gels with your ISP agreements and how systems would talk to each other if at all depending upon SLA, Performance & Visibility/Reporting requirements both may agree upon

5. Policy Stitching - This is one of the hardest part to get across and more so in a multi-vendor environment. As mentioned earlier - beside the fact that most QOS books and vendor QOS courses don't cover much details around platform specifics and they just assume that one would figure it out, the things gets pretty complicated pretty quickly the moment you know that the QOS depends on:

- Platform and Specific Model you are using
- NOS version
- ASIC Architecture (ASIC Pipeline, Buffer, Memory type & speed, Over subscription, Queue/Dequeue algorithm etc.)
- Chassis specifics (in case you are using one as opposed to fixed form factor) - example VOQ, Switch Fabric, Fabric Generation, Fabric Modules Count etc...
- Supervisor Engine & Architecture beside its generation, CEF vs. dCEF kind of implementation specifics 
- Policy Framework supported by NOS - example hierarchical QOS, support for sub-interfaces/Logical interfaces, how policy aggregation works and in which direction etc.

6. Modern App. Architectures - Since these days some of the new buzz words in application space are Cloud Native Apps, Micro Services, Containers & Kubernetes etc. One might wonder how he/she would go about planning QOS for such environment which are highly dynamic in nature with complex topologies, both short & long lived flows with mix of interaction surfaces with other systems and tools such as distributed tracing to feedback into your QOS Mgmt. tool.

7. Complexity Induced by Networks - There are some very common network choices that every network architect makes at some point which further complicates the QOS implementation. These are perhaps some of those complexities which must exist in order to deliver the desired outcomes and are least avoidable such as:

- MC-LAG aka Port-Channels/Ether-Channels/Bundle Interfaces
- Multi-tenant Networks (Remember you only got few queues in HW)
- Dynamic Network Traffic Patterns (Even more so with TE Controllers) during stable conditions vs. failure conditions
- How TC gets implemented by a given vendor in given platform & NOS
- Inflated throughput & performance numbers by vendors (very common)
- Different SP QOS Models (Customer Facing vs. Core Facing) as they usually have no more than 3 bits or 8 classes to play around beside allocation models
- QOS in Dual Stack Networks vs. IPv6 only Networks
- Some nerd knobs such as QPPB
- Impact of Physical & Logical topology

Hope you find this helpful and lets continue with this in Part-2.

HTH...

A Network Artist ðŸŽ¨

Monday, October 18, 2021

Facebook Down Event - The dilemma of a CTO, Black Swans & Fallacies of IBN


While Facebook just seem to have published somewhat a lengthysh version of root cause analysis (https://lnkd.in/guptSB3u) for public about their recent worldwide network outage that made - Facebook, WhatsApp & Instagram completely cut out from internet, it must have raised some concerns in the worldwide CTO and CIO community.


Since historically they have been told that and what pretty much every vendor in networking industry is preaching about in terms of different ways & methods (Systems, People & Processes) to avoid such circumstances are bright & magical ideas such as:

- Automation & Orchestration
- Intent Based Networking (IBN)
- Software Defined Networking (SDN)
- Centralized Controllers
- Data Models
- Automated Test & Deployment Pipeline with Unit Tests (aka CICD)
- Reliability & Resiliency Engineering
- AI/ML OPS
- Network Design Principles (Hierarchy, Swim Lanes, Segmentation etc.)
- Streaming Telemetry
- Observability Tools
- Bright Engineers
- Testbed Equipment
- Network Modelling & Simulation Tools with Formal Verification
- Rigorous platforms testing (HW/SW)
- Single Source of Truth
- Chaos Engineering
- BCP Plan
- Correlation Tools & what not

But assuming if you go via this checklist, Facebook would probably have all checks against all these items and so would be any of the FAANG company at this stage.

" So assuming you are a CTO or CIO, what would you suggest as possible next steps to your CEO & board if you have been called up this week for a meeting to discuss about how do we ensure such events don't happen in our network ? "

So lets park the above question for a while and move to what reactions we have seen so far.

1. The usual suspect is, bad things happens and everything breaks at some point, focus on RCA...move on and ensure it doesn't happen again

2. Network Architects favorite answer.... " it depends "

3. Was it a People or Process issue ?

4. The conspiracy theory that FB was under a Cyber attack which they don't want to disclose

5. Blame BGP (the easy suspect) … interestingly we got 10000+ new BGP experts on twitter and LinkedIn overnight :) beside the fact that 99% of them hardly understand the BGP details since none of them looked at the problem from perspectives of "unintended consequences", "ripple effect", "interaction surfaces", "failure domains", " & so forth beside all the pointers list I shared above. So let's say blaming BGP was an easy pick for the "ghost" network engineers. Beside the fact that RCA published by Facebook doesn't cover any technical details either.

6. "The Black Swans" - This is an interesting one and less talked about fact in case of this outage. While some may claim this was just one of those black swan events, I personally seriously doubt that and more so in the absence of a detailed RCA.

The answers lies in Systems thinking, Anti fragile networks, deep expertise, people's attitude towards excellence, and willingness to collaborate under uncertainty.

Further Readings:
















HTH...

A Network Artist ðŸŽ¨

Why Intent Based Networking (IBN) will Not Save Your Network Anytime Soon ?

 



In every few years, your favorite networking equipment vendor/provider/manufacturer comes to you with something new, suggesting this is the thing you need to solve all your networking problems in order to deliver the " Business Outcomes " that your CxO team is most desperate to achieve since very next second after the " Big Bang ".

Over the years both Networking Industry & Networking OEMs keep coming with new ideas 💡 - Separation of Planes, Centralized life cycle management, BGP as the mothership protocol by fitting every possible thing within that, Separating Policy from - Topology Information & Reachability Information, The magical 3 letters word " SDN ", Policy Based Networking, Application Aware Networking, Automated remediation and last but not least " The Magical Intent Based Networking aka IBN " beside AI Ops and what not.

But, You often forgot to involve qualified business people in defining those modern networking standards and terms, and for whatever reasons neither side ever took interest to step into the other side of the territory to get its basics right or to discuss what's really needed.

So the choice you are left with at the movement is - either you start taking those small steps or wait for IBN to fail (Since SDN and Open Flow have been declared dead already) on delivering those " Business Outcomes " and the cycle keeps on repeating endlessly.

And in case you doubt this, remember:

1. The intent based networking draft doesn't define or talk about " Business Outcomes " or even " Business Intent ".

2. Crazy Network Automation nerds should go and read millions of lines of code that a lady wrote for Apollo 11, which A. Tells you there is fundamentally nothing new in doing that and B. You can only automate better if You truly know the science behind stuff that you are trying to automate but more importantly have been able to frame the problem well enough. But let's circle back on those details and why often network automation fails to deliver true " Business Outcomes " for later post.

3. Most Network Engineers are not on party list when it comes to CxO meetings and most CxOs are either not invited for Network Strategy and Planning Discussions. Beside some exceptions where either party is there to have a cup of good coffee ☕

So you either start living with those ideas and keep chasing the unicorns, or you can take a step back and start thinking how to really get it right...or let's have a chat 😉 since you must remember that living under a rock is a choice...

HTH...

A Network Artist ðŸŽ¨

Sunday, August 8, 2021

Where & Why did We Lost SDN ?

 



In theory, the core objectives of Software Defined Networking (SDN) & Intent Based Networking (IBN) should have been (As an Integrated Package):

1. Introduce System Thinking

2. Simplified Life Cycle Mgmt.

3. Introduce New Abstraction Layers to Help Businesses People Have a Realtime/Historical View of Business Performance Metrics & Insights

4. Seamless Integration between Technology Metrics & Businesses Metrics to Fill the Gap between The Two

5. Reduced Complexity & Running Cost

6. Flexible Consumption Models

7. Embedded Security

8. Flexibility to Build Innovative Solutions on Top of Given Platform that Meets My Specific Environmental Needs

9. Gel Well with Existing & Legacy Environment and Tools

10. Allow Me to Run Dry Runs, Simulations and Offer Me Insights & Recommendation to Keep Improving + Troubleshooting

11. Reduce my Risk Profile (Technically & Financially)

But...What We Ended Up with:

0. Fancy GUIs Built Around Object Oriented Programmability Principles, You Just Need 10 Clicks to Create VLAN & Map it to a Port 😇

1. Separation of Control Plane & Data Plane

2. Network Virtualization

3. Network Function Virtualization & White Boxes

4. Human Driven Programmability & Automation

5. Controllers

6. Tons of Overlay Protocols

7. Proprietary Solutions


HTH...

A Network Artist ðŸŽ¨