Showing posts with label Network Designing. Show all posts
Showing posts with label Network Designing. Show all posts

Monday, September 11, 2023

The Data Centre Networking (DCN) Ref. Frame - Expanding MONAF Part 1

Your Data Centre Network (DCN) essentially is a stretch that cuts across SaaS, PaaS, Public Clouds, Hybrid Clouds & the Modern Day Multi-Cloud.

If You look underneath, it's essentially a set of technical capabilities, combined together to deliver business outcomes.

But what should be Your starting point to have a well structured conversation with business as well as the technical folks around DCN ?

Well, if we expand the DCN frame under the MONAF umbrella, it exactly does that to offer You are a very detailed and a firm structure by segmenting those capabilities into well defined and structured layered approach. (Essentially following the famous MECE consulting approach).

Let me know if there is any questions that comes in your mind around DCN that this frame couldn't fit into one of it's layers.


HTH...

A Tech Artist ðŸŽ¨

Monday, April 10, 2023

The Mother of All Network Architecture Frameworks - MONAF

 


In case You have been into the IP Networking Industry for long enough, You would have probably come across:

  1. The Holy "OSI Model"
  2. ISO Model
  3. TCP/IP Model
  4. RINA Model by John Day
  5. Hour Glass Model by Micah Beck
  6. SOS Model
  7. OODA Loop for Security

While all of these models are helpful in some form and shape, They often help us answer some of the very few questions we generally have as Network Architects, Designers & Engineers but often lacks:

- A bird eye view of Network Architecture as a whole

- A layered approach to go deeper into details & nuts and bolts while still allowing abstractions to work to meet requirements of different audience

- Fill the gap between understanding of Network among - Business, Architects, Designers, Engineers

- Address both functional & Non-Functional requirements

- Work as a guiding compass to address - "Known Knowns", "Known Unknowns", "Unknown Unknowns"

- Stitching together all of the information in a more tangible & informed way

- Last but not least, helps understand your specific constraints beside looking at everything across its life cycle

If you have been there and often relying on your head to hold so much of information but yet unable to put it into a structural frame, You might just want to give this a try.

In the subsequent posts we will uncover different layers and that too 5 layers down in terms of details (From 10,000 feet to 100 feet journey)

HTH...

A Tech Artist ðŸŽ¨

Thursday, January 5, 2023

An Architectural Perspective on Hierarchy In IP Networks - A Complex Puzzle Comprising Protocols, Topologies, Addressing & Systems (A Short Post)

 


If you ever bump into a Network Design book, course, blog or a webinar - most likely you are going to get introduced to this interesting & an important concept of " Network Hierarchy aka Hierarchical Networks Design Principle ".

Now depending upon which study materials and authors you follow, you would likely to come across different view points in terms of it's needs, pros & cons. Which in general not only contributes into more confusion among audience but also when I speak around with experiences Network Architects & Design Engineers, I often find them:

1. Having different interpretations of this concept and different view points

2. Considering this to be a very theoretical concept which you are likely to encounter in most Architecture & Design books but don't know about:

A. How to practice it (By applying theory to practice) 

B. How to measure it 

C. Missing the deep understanding of the topic at hand beside failing to understand its tradeoffs

So Idea behind this post is to offer you some architectural decision pointers to think through the problem statement and break it down into few tangible pieces by following another important network design principle " Separate the Complexity from the Complexity - Russ White " beside examining the rule 8 from RFC-1925

One of the model I personally always find handy is the SOS model from Russ White and you can use it too as a good ref. point.


So here is the quick list for you to think through in a more pragmatic manner:

1. What problems are you really trying to solve by introducing hierarchy into the Network (Go beyond theory) ?

2. Is it always possible to follow hierarchy? Specially in brown fields or during transitions (Think of old gear with still some lifetime left, mergers )

3. What are the downsides of introducing hierarchy ? (What harm it can cause and tradeoffs such us downgrade Agility, Flexibility, Organic Growth etc.)

4. Difference between Hierarchy vs. Symmetry vs. Modularity vs. Abstractions 

5. Different types of hierarchy/ layered approach to it ( physical level hierarchy,  logical level hierarchy,  hierarchy in addressing scheme, Protocol Level Hierarchy (ISIS Levels & Addressing ?) and so forth)

6. What data points you have in place to test your hypothesis to measure its impact on network

7. How these concepts are applied to different network environments - Enterprises (Campus <Wired and Wireless>, WAN/SDWAN, DC) vs Teclos vs CDNs vs Cloud Providers vs Web Scales vs Within public cloud virtual DC + Controller vs. Controller Less Architectures

8. Impact of introducing hierarchy on Visibility, Reporting and Performance mgmt. of the network

9. Impact on hierarchy on information hiding <reachability information> vs. topological information hiding (Aggregation vs. Summarization)

10. How all these choices will flow into your equipment sizing and potentially have an impact on your decision process

11. How will you apply all these concepts in a IPv10 network (IPv4 + IPv6 aka Dual Stack)

12. And if you are still brave enough :) , read through the further readings list to get to the bottom of this rat hole

Further Readings:

P-FatTree: A Multi-channel Datacenter Network Topology

Enabling Wide-spread Communications on Optical Fabric with MegaSwitch

Abstraction in Networks with Russ White

Hierarchical Network Design Overview

Engineer Versus Complexity

Optimal Routing Design

Navigating Network Complexity

Network Topologies

Five Number Summary for Network Topologies

Scaling MPLS Networks

The Side Effects Of Route Summarization

Avoid Summarization in Leaf-and-Spine Fabrics

Valley-Free Routing

Intra-Spine Links in Leaf-and-Spine Fabrics

Nonblocking versus Noncontending

Hierarchical IP Address Design and Summarization

Hierarchical IPv4 Framework

Fabric versus Network: What’s the Difference?

Liskov Substitution and Modularity in Network Design

Dragonfly+: Low Cost Topology for Scaling Datacenters

Reliability Basics- Part1

Network Centrality and Robustness

Swimlanes, Read-Write Transactions and Session State

Fifty Shades of High Availability

HTH...

A Network Artist ðŸŽ¨

Wednesday, January 4, 2023

Overlay Networks & Protocols Tradeoffs - Aka SDN aka IBN aka Magic aka Silver Bullet

 


A long time ago I wrote a short article on what really went wrong with SDN, now a few years later the topic still pops up in a conversation with the great Ivan & he acknowledged my list of Tradeoffs (things to watch out for carefully) related to overlay networks and protocols which seems to be de-facto standard for most modern Network solutions that we see around in Enterprises & Telcos.

  • Impact of overlay networks on visibility, reporting and performance management
  • Additional control plane that would result in additional abstraction layers and interaction surfaces and hence cascading effect in many situations
  • Impact on troubleshooting: how many solutions do we see in the market that can correlate underlay and overlay problems?
  • When it comes to sizing equipment in terms of control plane or data plane, it poses a new level of complexity an architect would need to deal with and in most cases vendors themselves won’t be able to offer much help in general rather than just asking you to believe in their words
  • I see lot of VXLAN and EVPN preachers, but let’s agree that mapping VLAN to VXLAN on 1:1 basis tells me you don’t know your stuff and believe too much in vendor marketing
  • EBGP underlay with IBGP overlay…man we can do better
  • Stitching two EVPN DCs with MPLS and SR: most of the implementations that I have seen were too complex and too fragile and thus results in a complex “policy.”

Further Readings:

Disjoint Path Routing and LP

HTH...

A Network Artist ðŸŽ¨

Monday, October 31, 2022

Marrying SASE (SDWAN) with 5G - The Marketing, The Myths & The Fallacies & How to Get it Right (An Architectural Perspective)

When almost 3 years ago I wrote about why having an inbuilt LTE interface inside a SD-WAN device doesn't really matter, the 5G thingy was still relatively new.

During a recent Enterprise Architecture Consulting engagement, I was asked by one of my client if they should really care about 5G and 5G interfaces on the variety of SD-WAN platforms that were pitched to them by different Systems Integrators (SIs) & MSP (Managed Service Providers)/Telcos.

So let's start with a simple question - "What problem we are trying to solve?"

In general you will see a few types of customers in SASE/SD-WAN market :

1. Which have Technical/Solutions architects those are completely sold on vendors marketing (50% of the crowd)

2. Those who wants to jump on the bandwagon due to fear of being left behind in the similar industry or by the competition (25% of the crowd)

3. Those who always are either too excited by technology or have a lot of money to throw onto the problem (the next 20%)

4. Those who can really map business capabilities to technology capabilities (the rare and the last 5%)




In general, you would often find a few ways the 5G gets included into the solution by solution providers such as :

Design 1 - You have a site (mid/large size) which either has got hybrid connectivity (MPLS + Internet) or 2 x Internet connects, while keeping 5G cellular as a last resort backup link in an event of a total failure.

Design 2 - A small site that usually runs on a single internet link and keeping 5G as back for last resort.

Design 3 - 5G as backup of last resort onto your DC WAN edge

Now in general it doesn't look like a bad idea to have a 5G interface but here are few of the important considerations:

- In general your DCs/COLOs and Campus networks would be setting on racks behind thick physical building structures (remember your Wi-Fi coverage problems even while your APs are sitting inside), so very often you would expect coverage issues. Now one of the argument here might be that I can install an external antenna of some sort on top of building or floor and run a fiber cable connection from there. Fair...but:

> Now you got to get approval for installations, cable runs, have administrative processes and safety processes in place and what not (What if the lightening hit the antenna?).

- To my understanding, most SASE/SD-WAN solutions don't offer any visual monitoring, reporting & troubleshooting tools for - Checking 5G signal strength, 5G interface troubleshooting, Dummy traffic probes etc.

- Interestingly enough now you need even a more complex traffic distribution, traffic prioritization, traffic failover, traffic desired SLA/performance metrics and other set of policies into the mix. And even if you end up doing that successfully, how you are going to document it for the operations ?, Is your EMS/NMS equipped with such capabilities ? 

- From the network architecture perspective, you just added an another layer of complexity. Assuming this 5G interface is an HW module, you got to now deal with: New stack of software and protocols within your fancy WAN edge device (3GPP standards), New interaction surfaces, New potential grey failures.

- You just end up adding the more state into the network (State, Surface & Optimization tradeoffs

- You need to now have life cycle mgmt. in place for your 5G interface (HW/SW Upgrades, Monitoring, Management, Refresh etc.) beside that fact that the 5G specifications often vary country to country (even from Telco to Telco) and now you need to keep track of Data Plans, Data Usage, Availability & Performance Mgmt., Cost Mgmt. and what not. (Remember your data plans are pretty limited in general?)...imagine to solve these problems at a global scale deployment dealing with different MSP.

> How to you move to 6G if it comes out in the next few years ?

Now after all these interesting questions, we may still ask:

1. What are the better alternatives today ?

2. Where 5G might still make sense ?

Answering the first question, IMHO I would still recommend you to opt for a broadband connection and avoid 5G unless you have a very particular problem or scenario because:

- With broadband you are still dealing with Ethernet connection between your CPE and broadband router/device which is a pretty familiar connectivity model and protocol stack to deal with.

- In general your broadband data plans are much bigger 

- You don't need to deal with another MSP for service management perspective as in general your ISP would have a common portal to give you a view of all MPLS, Enterprise Class Internet and Broadband based Internet connections. (Also mind that historically your SPs were building two parallel networks for ISP (Internet Service Provider) and MSP (Mobile Service Provider) business units, though those are converging now more and more)

- With everyone hit by pandemic that accelerated WFH culture, in both developing and developed countries you would expect to have broadband being available very easily and at the affordable prices. The only places you might still face availability issues are Tier-3 cities and so forth. But for most part that is not a technology problem but your SPs wanting you to stick with cellular connectivity to drive more profits (So its an intent problem rather)

- My own research suggests that in many developed countries getting a broadband will cost you far less compare to the enterprise class 4G/LTE/5G/Private 5G connection

Answering the second question, here are few use cases for your consideration:

- 5G as part of your SD-WAN transition plan (Legacy to SD-WAN)

- Your last mile hybrid or other wired connectivity model still runs on same fiber or shared media (true path diversity problem)

- One or both of you last miles are on Microwave (Should be pretty rare now)

- 5G as a OOB (Out of Band) management option

- Movable workplace and offices (Eg. Marketing/Sales & Promotion offices or moving semi-trailer trucks often used in variety of businesses)

- Specific operating conditions such as in Oil/Gas & Mining industries

- Edge compute/Cloud

- IoT Platforms

HTH...

A Network Artist ðŸŽ¨

Further Readings:

SD-WAN Leads to $96,000 4G/LTE Bill

Improve Your Home Internet Performance Using CoDel

More Bandwidth Doesn’t Matter (much)

Are Networks Really Complex ?

Enterprise QOS Design & Deployment - Good, Bad Or Ugly ?

Focus on Your Business, Not Fancy Technologies

Are You Solving the Right Problem?

This Is What Makes Networking So Complex

Are Business Needs Just Excuses for Vendor Shenanigans?

The Three Paths of Enterprise IT

SDN Will Not Solve Real-Life Enterprise Problems

Why Intent Based Networking (IBN) will Not Save Your Network Anytime Soon ?

Complexity and the Thin Waist

It’s Most Complicated than You Think

Details and Complexity

Monday, May 16, 2022

How Many Routes My ASIC Can Hold ? - A Short Post

 


One of thing that always keep surprising me over the years is why 90% of pre-sales people in tech can't size the equipment correctly or don't have a solid logic to make those decisions, hence 99% of equipment are oversized to play safe and businesses have no clue that they are paying so much more. 

Obviously part of the problem is they have never come across formal training, course or book talking about it, even the most popular design and architecture books fails to address this or offer a comprehensive view/approach.

So instead of treating it as a science, most will resort to tactics.

In late 2021 I wrote this brief article around skills a network engineer should pick on in his/her early career. Where I suggested to be at least familiar with basic understanding of both "Router Architecture" & "ASIC Architecture". Now obviously the depth is always subjective to:

1. How many details I need to know for my current role & responsibilities in order to get things right

2. The amount of details and depth I need to know for potential future roles that you might be targeting 

3. Your personal curiosity & interest 

4. If you are really into Architecture & Design, You got to have fair & intermediate level of understanding of these topics at minimal

5. If you into Pre-sales, You often got to deal with sizing & performance for a given set of equipment as part of your deliverables requested by client in the form of RFP or RFI. Remember those data sheets you often have to refer to claiming IPv4 or IPv6 prefixes numbers a platform can hold/support ?

6. You might have to do platform testing at some point as part time or full time job including you may land yourself into a COE (Centre of Excellence) of your organization or may end up into a "Platform/Service Product Management" role.

Now you must know that often those specifics and details are hidden and never publicly shared/offered by most of ASIC/Platform vendors. You often got to be a premium customer and sign-off tons of NDAs to get those details to some extent and more importantly you got to be very specific around what exactly you are looking for since asking for data in abstraction would often result into tons of non-specific information thrown on you by your fav. ASIC vendor.




Assuming by now you have some more clarity in terms of why you need to know all these details as a Network engineer depending upon where you are and where you plan to end up, lets circle back to original topic for today which is "How many IPv4 (could be IPv6) prefix my device support in reality?"

Which leads us to a simple question - "What are the different variables I am dealing with when trying to figure how much routes my platform can really hold?"

While a simple answer you would often hear would be "it depends" or someone may point your to RFC-1925 rule 8 "It is more complicated than you think"

So let's try to list some of them in this series Part - 1

  •  ASIC Architecture 
    • ASIC pipeline
    • Memory architecture
    • Memory Carving/allocation to different features & functions (HW/SW)
    • How the information is queued & de-queued 
    • API details (Type of API, API interface, Information flow etc.)
    • Routing Vs. Switching ASIC
    • Hierarchical vs. A Flat FIB
  •  NOS Architecture  
    • How NOS is programming the FIB
    • Prefix Length 
    • Contiguous Vs. Dis-contiguous Prefixes
    • Sorting Algorithm & Data Structures 
    • Device Profiles/Resource Allocations by NOS 
    • NOS Scheduler
    • ECMP, UCMP, FRR
  •  Platform Architecture
    •  Line Card Architecture
    •  Back Plane Architecture 

Further Readings:

ASIC

ASICs for Network Engineers




A Brief History of Router Architecture






Anatomy of Core Network Elements

SONiC: Open Source NOS in Data Cente

Cisco - Configuring SDM Resource Allocation Templates

Adjacency Matrix, Adjacency List, Priority Queue Implementation 

Juniper Networks Routing ASIC Strategy

Cisco 8000 Series - Under the Hood



Sizing the Buffer

Sizing Router Buffers - Small is the New Big...

Embedded Hardware for Processing AI at the Edge: GPU, VPU, FPGA, and ASIC Explained

ASICs vs. Net Processors: Understanding the True Costs

P4 - Programming Protocol-Independent

Open Flow Specifications - Remember how Open Flow Originally planned to program the ASIC directly using Open Flow Controller ?

How Routers Really Work - A Webinar from Russ White under O'Reilly Subscription

Networking Hardware/Software Disaggregation in 2022

Select the Best Switching ASIC For the Job

Data Center Switching ASICs Tradeoffs

FIB Compression

Switching Hardware Series - Part 1 , Part 2 & Part 3

Juniper MX10000 LC480 Deepdive

BGP RIB Sharding

Using Trio -- Juniper Networks' Programmable Chipset -- for Emerging In-Network

Optimizing Power Consumption in High-End Routers

Striking a Balance: Exploring Fairness in Buffer Allocation and Packet Scheduling

Making 35 000 000 IP lookup operations per second with Patricia tree

Optimizing Power Consumption in High-End Routers

Saving Energy on PTX with PFE Power Off

ACX7000 L2 MAC Scale and Learning Rate

FIB Compression in Juniper Routers

PTX10001-36MR FIB Install Rate

Express 4 Filters - Foundation

Large Language Models — the hardware connection

Classification TCAM with Cisco CloudScale ASICs for Nexus 9000 Series Switches White Paper

Chiplets - The Inevitable Transition

[Podcast] The chips are down: Moore’s Law coming to an end


HTH...

A Network Artist ðŸŽ¨

Wednesday, December 22, 2021

The QOS Fallacies & Failures in a Modern Hybrid IT World - Part 1 of 2

 


When I wrote about QOS the last time around 7 years ago, I must say I had high hopes from SDN & IBN as both of the paradigms were still evolving. Meanwhile there have been some unsuccessful attempt to automate QOS by throwing some sort of controllers into the mix by few vendors beside some others claiming they can solve this problem with the mighty IBN.

As we are about to move into 2022, many still wonder if QOS makes any sense at all in the context of modern networking ?

In order to find the answers, let's break the problem into two parts:

1. Why QOS has been so unsuccessful historically
2. What are our options moving forward

So let's focus on point 1 to begin with.

1. How do we get started ? - Interestingly enough over a dozen books have been written on QOS over the last 2 decades or so in the context of IP networking which mostly talks about details such as congestion management vs. congestion avoidance and so forth. But very few of them actually jumps into platform specifics in terms of capabilities and dependencies (both HW & SW). More interestingly I personally haven't come across a single QOS book myself yet which gives you any practical advice or framework/methodology around how to gather technical requirements in reference to Applications in order to plan and craft a QOS policy. So often I have seen people struggling to come up with one and given most QOS deployments are tactical rather than strategic, people often have time constraints to come up with a one in a short time. That's why many times people end up coping some references from recommended design guides etc. which hardly works in real life (Unless you were too lucky !).

2. Benchmarking & Capacity Mgmt. - Most small, medium & even couple of the large enterprises that I have worked with including Telcos don't seem to have both of these as mature practices in place. Benchmarking is though one of the key exercises you need get through to craft a good QOS policy beside being a necessity in any of a mature Capacity mgmt. framework/practice. The other problem you may likely to run here is that in order to do effective benchmarking & capacity mgmt. you need to invest into additional visibility & performance mgmt. tools to gather the required details which are usually quite expensive beside that fact that you need to train your team on tools and required operating skills (for example statistical analysis, Time Series, Sampling details etc.). Certain times you are likely to run into the problem where in the given tool may not be able to offer you reporting/data that you need natively which means you are always dependent on tool vendor about if their product road-map is aligned to your priorities and timelines. And be careful about if the tool allows you to run custom reports or exports required data in format you need. So if your capacity mgmt. is still on excel sheets, you know where you are heading.

3. Measuring latency incorrectly - This is perhaps more common than you might have thought beside that fact that most QOS books don't offer any practical advice here too and details around how latency needs to broken down across the spectrum. Can you tell me the breakup of end to end latency (Server - Client) and that too on hop by hop basis ?

4. QOS Lifecycle Mgmt. - While this area has improved a bit when it comes to modern networking gear, assuming majority of equipment still out there are old ones which doesn't offer much when it comes to QOS lifecycle mgmt. that includes Plan, Design & Implementation. But more importantly what they lack are capabilities such as QOS monitoring & reporting in real time beside the correlation with network health & events. After all you don't want to hire someone today to type couple of show commands in every few minutes and running the scripts won't be that helpful either for most part when you are dealing with scale. Again there are couple of commercial and open source tools available but its an exercise which takes time and resources beside that fact you should know exactly what you are looking for in which scenario. BTW...will that resource be from planning team, tools team or ops team is what I leave for you to figure out in real life. :)

Also with the rise of modern solutions which are mostly built around the magical controllers and overlays, you must think about QOS FCAPS capabilities in such environment carefully. For example while from overlay protocol perspective everything might be just a single hop away, the packet eventually still gets passed through the physical world in underlay. So its importantly to find out early if QOS policies will be:

- Static or Dynamic in nature (Given you are using controller of some sort)
- Correlation of QOS statistics between underlay & overlay
- How policy gets propagated 
- How controller interacts with other systems and policies for introducing dynamic behavior 
- Does the system allows Time Based QOS policies (interestingly enough most don't yet)
- Dummy policy dry run capabilities if supported 
- How your QOS policy gels with your ISP agreements and how systems would talk to each other if at all depending upon SLA, Performance & Visibility/Reporting requirements both may agree upon

5. Policy Stitching - This is one of the hardest part to get across and more so in a multi-vendor environment. As mentioned earlier - beside the fact that most QOS books and vendor QOS courses don't cover much details around platform specifics and they just assume that one would figure it out, the things gets pretty complicated pretty quickly the moment you know that the QOS depends on:

- Platform and Specific Model you are using
- NOS version
- ASIC Architecture (ASIC Pipeline, Buffer, Memory type & speed, Over subscription, Queue/Dequeue algorithm etc.)
- Chassis specifics (in case you are using one as opposed to fixed form factor) - example VOQ, Switch Fabric, Fabric Generation, Fabric Modules Count etc...
- Supervisor Engine & Architecture beside its generation, CEF vs. dCEF kind of implementation specifics 
- Policy Framework supported by NOS - example hierarchical QOS, support for sub-interfaces/Logical interfaces, how policy aggregation works and in which direction etc.

6. Modern App. Architectures - Since these days some of the new buzz words in application space are Cloud Native Apps, Micro Services, Containers & Kubernetes etc. One might wonder how he/she would go about planning QOS for such environment which are highly dynamic in nature with complex topologies, both short & long lived flows with mix of interaction surfaces with other systems and tools such as distributed tracing to feedback into your QOS Mgmt. tool.

7. Complexity Induced by Networks - There are some very common network choices that every network architect makes at some point which further complicates the QOS implementation. These are perhaps some of those complexities which must exist in order to deliver the desired outcomes and are least avoidable such as:

- MC-LAG aka Port-Channels/Ether-Channels/Bundle Interfaces
- Multi-tenant Networks (Remember you only got few queues in HW)
- Dynamic Network Traffic Patterns (Even more so with TE Controllers) during stable conditions vs. failure conditions
- How TC gets implemented by a given vendor in given platform & NOS
- Inflated throughput & performance numbers by vendors (very common)
- Different SP QOS Models (Customer Facing vs. Core Facing) as they usually have no more than 3 bits or 8 classes to play around beside allocation models
- QOS in Dual Stack Networks vs. IPv6 only Networks
- Some nerd knobs such as QPPB
- Impact of Physical & Logical topology

Hope you find this helpful and lets continue with this in Part-2.

HTH...

A Network Artist ðŸŽ¨

Wednesday, November 17, 2021

A Simple Routing Protocols Decomposition Model - Part 2 (Peering Mgmt.)

 


So in the first part of the series, we started with rather a simple decomposition model to get bit more insights into the routing protocol internals.

So let's continue the series by expanding on the very first layer in our model - Peering Management.

In any routing protocol before we get fancy in terms of which all features, functions and knobs to use, the very basic requirement is to peer with other devices into the network since eventually a routing protocol is nothing but a distributed database. Routing protocols are designed to convey different set of information to its peer devices such as:

- Topology Information

- Reachability Information

- Policy Information

The type of information that a given routing protocol would exchange with its peers would largely depend on the protocol itself (OSPF, IS-IS, BGP) & where that proposed routing protocol is used into the network (Campus, WAN, Metro-E, DC) as implementation specifics do change.





[ Click on Image to Enlarge ]

As you may notice, there are lot of things working behind the scenes when it comes to peering in a routing protocol context. But don't get carried away by looking at the complexity. Once you look closely, all of these pieces kind of makes sense.

So let's start with the top row, reading it from left to right. 

Self Identity - Before the routing protocol determines whom it needs to communicate with and what information needs to be exchanged, it must find its own identity first. The most common way to give an identity to the routing protocol instance itself is assigning it a router id (RID). The RID can be configured manually or it can be derived automatically depending upon the platform and NOS

In real life assuming network virtualization is much more common today, you are allowed to configure unique RID for each routing protocol as well as a unique RID for each instance/process under same protocol. 

Protocol Addressing - Once the routing protocol is able to define its identify with a RID, the next step is to understand it's addressing. The addressing serves many purposes but the most basic one is to provide location services from the overall network view standpoint. Though protocols such as LISP was an attempt to separate device identity from device location, due to limited use cases (such as mobility) and other problems it never really took off well.

Every routing protocol has it's own addressing scheme which may further have impact on its scaling and expected working behavior if not done correctly.

Participation - The next step for the routing protocol is to determine its participating interfaces on a given device and in certain cases the entire device itself. Depending upon the design you may run into some interesting challenges though.

Reliable Transport - Every routing protocol needs a reliable transport to be able to effectively communicate with its peer device. The reliability itself is an important aspect and while some routing protocols such as BGP rely upon existing TCP stack, others like OSPF uses IP protocol 89 & EIGRP choose its own transport protocol namely RTP.

Neighbor Discovery - In the next step the protocol must discover its peer/neighbor/adjacent device depending upon which routing protocol you are following. The protocol while ideally should keep a track of its neighbor and relationship state, this may or may not be implemented.

The neighbor could be configured manually or dynamically discovered, while the discovery phase itself may use unicast or multicast as a transport for reachability purpose to the next hop device. Also the reachability to the neighbor could be over layer 2 transport or layer 3 transport depending upon the protocol. IS-IS and many other IOT industrial protocols operate at layer 2 for example. In case of eBGP, the neighbor in fact may be multiple physical hops away.

Neighbor Identity - While we may have discovered our neighbor, it doesn't mean the neighbor itself is a an intended neighbor or a legitimate neighbor always. After all someone might want to spoof or sometimes we may end up discovering somebody completely un-intentionally. So sharing any information with an unexpected neighbor won't make any sense. To prevent this we have several measures which we can put in place such as Authentication, Validating neighbor's identity (remember they also have RID), Validating neighbor based on IP Packet's TTL value etc. such as in OSPFBGP. The modern day solutions such as SD-WAN usually uses RPKI over TLS/DTLS channel.

Establish Session - Once the neighbor is discovered and validated, we finally establish a session with it. Depending upon the protocol, we might have a single session vs. multiple sessions going on. A simple example would be networks running IPv4 & IPv6 at the same time under single routing protocol instance. While some implementations exchange information related to both IPv4 and IPv6 over a single session, some may do it over a separate dedicated session for each of them. Long time ago there was an attempt to run multi session bgp for MTR (Multi Topology Routing) for network virtualization use cases.

Capabilities Exchange -  This is an another interesting step where the routing protocol running on separate devices exchange capabilities with each other to find the lowest common denominator. For example we know BGP is more like an application that runs on top of TCP as opposed to a pure layer 3 routing protocol which only carries routes. BGP though can carry layer 3 routing information and in case with most vendors it's been the default behavior, BGP does allow us to carry many other set of information depending upon the use case in the form of AFI/SAFI which is essentially an encoding format. For example BGP can carry MAC Addresses information under Layer 2 VPN EVPN address family when enabled. Though with BGP, you got to be cautious about enabling a new capability/address family in production network as highlighted here.

Establish Adjacency - The protocol reach this far and finally the peers are ready to exchange required set of information needed to populate RIB & other details such as topology graph. An interesting example in a routing protocol context would be (In case you are wondering in which scenario two devices would be neighbors but not adjacent) OSPF.

Messages Exchange - Finally we [assuming by now you are thinking like a routing protocol :) ] reach this stage where in we finally start exchanging information through messages. The messages needs to be sent reliably, keeping track of to understand which one to prefer in case of being received from multiple sources, acknowledged and so forth besides how to queue and dequeue them and at what intervals those should be sent out vs. being hold back for a while to pack multiple events together for optimization and getting the latest information being sent out.

Further Readings:

Cisco IP Routing: Packet Forwarding and Intra-domain Routing Protocols: Packet Forwarding and Intra-domain Routing Protocols 

Network Routing: Algorithms, Protocols, and Architectures

Network Algorithmics,: An Interdisciplinary Approach to Designing Fast Networked Devices

Inside Cisco Ios Software Architecture

HTH...

A Network Artist ðŸŽ¨