Showing posts with label QOS. Show all posts
Showing posts with label QOS. Show all posts

Wednesday, December 22, 2021

The QOS Fallacies & Failures in a Modern Hybrid IT World - Part 1 of 2

 


When I wrote about QOS the last time around 7 years ago, I must say I had high hopes from SDN & IBN as both of the paradigms were still evolving. Meanwhile there have been some unsuccessful attempt to automate QOS by throwing some sort of controllers into the mix by few vendors beside some others claiming they can solve this problem with the mighty IBN.

As we are about to move into 2022, many still wonder if QOS makes any sense at all in the context of modern networking ?

In order to find the answers, let's break the problem into two parts:

1. Why QOS has been so unsuccessful historically
2. What are our options moving forward

So let's focus on point 1 to begin with.

1. How do we get started ? - Interestingly enough over a dozen books have been written on QOS over the last 2 decades or so in the context of IP networking which mostly talks about details such as congestion management vs. congestion avoidance and so forth. But very few of them actually jumps into platform specifics in terms of capabilities and dependencies (both HW & SW). More interestingly I personally haven't come across a single QOS book myself yet which gives you any practical advice or framework/methodology around how to gather technical requirements in reference to Applications in order to plan and craft a QOS policy. So often I have seen people struggling to come up with one and given most QOS deployments are tactical rather than strategic, people often have time constraints to come up with a one in a short time. That's why many times people end up coping some references from recommended design guides etc. which hardly works in real life (Unless you were too lucky !).

2. Benchmarking & Capacity Mgmt. - Most small, medium & even couple of the large enterprises that I have worked with including Telcos don't seem to have both of these as mature practices in place. Benchmarking is though one of the key exercises you need get through to craft a good QOS policy beside being a necessity in any of a mature Capacity mgmt. framework/practice. The other problem you may likely to run here is that in order to do effective benchmarking & capacity mgmt. you need to invest into additional visibility & performance mgmt. tools to gather the required details which are usually quite expensive beside that fact that you need to train your team on tools and required operating skills (for example statistical analysis, Time Series, Sampling details etc.). Certain times you are likely to run into the problem where in the given tool may not be able to offer you reporting/data that you need natively which means you are always dependent on tool vendor about if their product road-map is aligned to your priorities and timelines. And be careful about if the tool allows you to run custom reports or exports required data in format you need. So if your capacity mgmt. is still on excel sheets, you know where you are heading.

3. Measuring latency incorrectly - This is perhaps more common than you might have thought beside that fact that most QOS books don't offer any practical advice here too and details around how latency needs to broken down across the spectrum. Can you tell me the breakup of end to end latency (Server - Client) and that too on hop by hop basis ?

4. QOS Lifecycle Mgmt. - While this area has improved a bit when it comes to modern networking gear, assuming majority of equipment still out there are old ones which doesn't offer much when it comes to QOS lifecycle mgmt. that includes Plan, Design & Implementation. But more importantly what they lack are capabilities such as QOS monitoring & reporting in real time beside the correlation with network health & events. After all you don't want to hire someone today to type couple of show commands in every few minutes and running the scripts won't be that helpful either for most part when you are dealing with scale. Again there are couple of commercial and open source tools available but its an exercise which takes time and resources beside that fact you should know exactly what you are looking for in which scenario. BTW...will that resource be from planning team, tools team or ops team is what I leave for you to figure out in real life. :)

Also with the rise of modern solutions which are mostly built around the magical controllers and overlays, you must think about QOS FCAPS capabilities in such environment carefully. For example while from overlay protocol perspective everything might be just a single hop away, the packet eventually still gets passed through the physical world in underlay. So its importantly to find out early if QOS policies will be:

- Static or Dynamic in nature (Given you are using controller of some sort)
- Correlation of QOS statistics between underlay & overlay
- How policy gets propagated 
- How controller interacts with other systems and policies for introducing dynamic behavior 
- Does the system allows Time Based QOS policies (interestingly enough most don't yet)
- Dummy policy dry run capabilities if supported 
- How your QOS policy gels with your ISP agreements and how systems would talk to each other if at all depending upon SLA, Performance & Visibility/Reporting requirements both may agree upon

5. Policy Stitching - This is one of the hardest part to get across and more so in a multi-vendor environment. As mentioned earlier - beside the fact that most QOS books and vendor QOS courses don't cover much details around platform specifics and they just assume that one would figure it out, the things gets pretty complicated pretty quickly the moment you know that the QOS depends on:

- Platform and Specific Model you are using
- NOS version
- ASIC Architecture (ASIC Pipeline, Buffer, Memory type & speed, Over subscription, Queue/Dequeue algorithm etc.)
- Chassis specifics (in case you are using one as opposed to fixed form factor) - example VOQ, Switch Fabric, Fabric Generation, Fabric Modules Count etc...
- Supervisor Engine & Architecture beside its generation, CEF vs. dCEF kind of implementation specifics 
- Policy Framework supported by NOS - example hierarchical QOS, support for sub-interfaces/Logical interfaces, how policy aggregation works and in which direction etc.

6. Modern App. Architectures - Since these days some of the new buzz words in application space are Cloud Native Apps, Micro Services, Containers & Kubernetes etc. One might wonder how he/she would go about planning QOS for such environment which are highly dynamic in nature with complex topologies, both short & long lived flows with mix of interaction surfaces with other systems and tools such as distributed tracing to feedback into your QOS Mgmt. tool.

7. Complexity Induced by Networks - There are some very common network choices that every network architect makes at some point which further complicates the QOS implementation. These are perhaps some of those complexities which must exist in order to deliver the desired outcomes and are least avoidable such as:

- MC-LAG aka Port-Channels/Ether-Channels/Bundle Interfaces
- Multi-tenant Networks (Remember you only got few queues in HW)
- Dynamic Network Traffic Patterns (Even more so with TE Controllers) during stable conditions vs. failure conditions
- How TC gets implemented by a given vendor in given platform & NOS
- Inflated throughput & performance numbers by vendors (very common)
- Different SP QOS Models (Customer Facing vs. Core Facing) as they usually have no more than 3 bits or 8 classes to play around beside allocation models
- QOS in Dual Stack Networks vs. IPv6 only Networks
- Some nerd knobs such as QPPB
- Impact of Physical & Logical topology

Hope you find this helpful and lets continue with this in Part-2.

HTH...

A Network Artist ðŸŽ¨

Monday, March 9, 2015

Enterprise QOS Design & Deployment - Good, Bad Or Ugly ? - The Business Side (Case Study)


For last couple of years I have been part of couple of Enterprise Level QOS Deployments. While some of them were tactical and others were completely strategic.

Now QOS is one of those topics in Network Industry which are considered to be highly complex and misunderstood at the same time. The part of the equation is there are lot of moving pieces that must fit together correctly in order to deploy QOS successfully in an Enterprise Environment.

At times I have seen Engineers just copy and paste QOS configurations from some other Enterprise they had access to in past and hoping that would solve the purpose. While in other cases people design QOS policies ensuring top priority for voice and video traffic in an Enterprise Network.

Now many Network Engineers while deploying QOS policies think that they should give highest priority to Voice and Video traffic in their Network. Now one of the common problem here is " Assumptions ".

As a Network Engineer you should always first observe the current state of Network Architecture, Hardware In Use, Documenting List of Critical Business Applications and Understanding their deployment along with Network requirements such as Transport (TCP vs UDP) (One to One Flow, One to Many Flow or Many to Many Flow) (Latency requirements) (Direct vs Redirected Access) (Realtime Vs Interactive) (Transport - LAN vs WAN, Wired vs Wireless) (SP Dependencies, SLA & Service Agreements) (MPLS Layer 2 vs Layer 3 VPNs) etc

The problem usually is that as Network Engineer we always try to understand and solve problems using technologies and tools. Now understanding technology and tools is definitely an important piece here but at the same time you should be able to convert a given business requirement into technical design and solution.

For example most books written around QOS would tell you to mark Voice with Highest Priority like EF and Video probably with AF41. While Voice and Video both are quite sensitive as traffic in a given network but does that also mean those are the most critical ones from Business Standpoint ?. Well if you think carefully that might not be the case in reality. For example an Enterprise client I have been recently working for was into News Paper Business. Now the Editorial Application they use to make news paper had a very tight Latency requirements end to end which was 40msec or less. While If we compare this with voice traffic latency guidelines which is 150msec or less we can certainly find the given requirements are very tight. At the same time this brings an interesting question from design standpoint which is " Shall we still give EF marking to Voice or Shall we use EF marking for Editorial Application ? ". Now if we think from business standpoint or talk to business to figure it out - The answer most likely is going to be that Editorial Application shall be given most priority. Now to make situation a little more complex they had an old Editorial application which was still in use while they were moving to new Editorial Application , so in that sense we now have two Editorial Applications with high latency sensitive requirements :). So we must take of these considerations while designing QOS policy.

Now would QOS policy alone solve the purpose now ?. Well as I mentioned the end to end latency requirements for given Editorial Applications were 40mses or less, which means you need to take a look at WAN Architecture of customer and see how this requirement can be met. Now the company had One Hub Locations in each region of India with approximately 50 remotes sites connecting to each hub site. All Hub locations were connected in partial mesh fashion.


Now the customer WAN comprises 2 MPLS L3 VPN service providers and tons of Point to Point WAN CKTs. Now at high level all looks good. At max we need to ensure that our MPLS VPN service providers are in agreement to accept our QOS markings and give our traffic proper treatment across SP core.

Now the twist here is that while most of remote locations had Cisco ISR G1 or G2 routers, the Core locations were using Cisco Metro Ethernet Switches as MPLS CE device. Now interestingly most Metro Ethernet switches don't support Layer 3 QOS for the purpose of Bandwidth based reservations under LLQ but only L2 QOS and MPLS QOS. So again from the traditional QOS deployment standpoint it could very well be a major pushback.

So as you can see the couple of technical reasons and less understanding or communication with Business makes QOS a complex and misunderstood topic.

The other problem I have seen in field is while Network Engineers try to make policies for things such as Bandwidth reservations under queuing , they don't do traffic pattern analysis properly to figure out correct bandwidth requirements. Ideally you should deploy tools such as NetFlow and let it run for couple of weeks and later analyse it to reach on conclusions for bandwidth usage vs reservation requirements.

So as you can see from this brief post on Non Technical Side of QOS, there are perhaps too many pieces involved to make a QOS deployment successful. While most people say that WAN is having highest potential in terms of SDN use, QOS is probably another key area where SDN and Automation have huge scope IMHO  [ Ever tried to deploy LAN QOS with couple of 6500s, 4500s, 3750, 2950, Nexus in a single setup ? :) ]

HTH...
Deepak Arora
Evil CCIE

Monday, June 25, 2012

IPexpert's Protocol Operations and Troubleshooting Series





In last couple of months, one of the most popular CCIE R&S training vendor in market - IP EXPERT released couple of books covering R&S blueprint technologies. Which I personally feel is a big effort. Since writing books not only takes time but also those should be user friendly as well. Otherwise there are always choices in terms of books. The best part about IPX books is that they not explain you technology in a very simple manner but also focus on configuration examples to make sure you really understand the concepts and also focus heavily on troubleshooting. Troubleshooting is I guess the part which most book vendors fail to cover or help you as reader to build step by step logical approach.


So far they have released four books under this series and I wish they will keep doing it for rest of important R&S technology domains.


http://www.ipexpert.com/Cisco/Troubleshooting-Series


HTH...
Deepak Arora
Evil CCIE

Friday, July 29, 2011

Rate Limit Calculator AKA CAR (Committed Access Rate)

Recently I have been asked for quick method to calculate " CAR Parameters also known as Rate Limit ". 





So don't confuse this CAR with our well known CAR :-)

Anyways... here is a great work done by "Brian" on Cisco learning Network site for your help... awesome work I would say.


BTW.... In modern days we have QOS tool called " policing " which is essentially modern way of doing CAR using MQC (Moduler QOS CLI).

HTH...
Deepak Arora

Tuesday, May 31, 2011

Back To Back Frame Relay - Old Days Stuff

So today I am gonna discuss some old school stuff called "Back To Back Frame-Relay". So the first thing coming in your mind would be "What Exactly is that ?" & Second thing will be "Why do we need this ?"



So coming to Idea of why do we need this. Actually in old days and galaxy far far away everyone was running P2P circuits. As you might be aware on those P2P links we use to run PPP. Though Cisco was cool enough and came up with HDLC.

But the problem was that there was no Native QOS Mechanism supported by PPP or HDLC by own. And of-course in those days we didn't have QOS based on MQC. 


So one day one evil engineer came up with idea "Alice - Hey Bob ! I know Frame-Relay supports a native QOS mechanism called FRTS(Frame-Relay Traffic Shaping) so why don't we use to apply QOS on our P2P links? "


But Frame-Relay on Point to Point Circuits ? How we gonna do that ? What we gonna name this thing ?

Hmmm... We call this Back To Back Frame Relay :-)


After this cool story now lets take a look at our simple topology and configuration we require.














 


We disabled Keep Alive in this implementation since there is no Frame-Relay Switch involved in topology(or in P2P links). Also on both sides we must use the same DLCI number.

So now you know that Back To Back FR was only method in existence in old days to implement QOS on P2P links for stuff like VoIP etc. All you gotta do is to create "map classes" to apply FRTS on these links.

HTH...
Deepak Arora



 

Monday, September 27, 2010

QOS Network Design Guide - Cisco

Today I found a great QOS resource from Cisco Website. The QOS Network Design Guide from Cisco.

The guide covers design guidelines for LAN QOS Vs WAN QOS Vs MPLS QOS etc.

A gr8 QOS Resource Indeed.


HTH...
Deepak Arora

Friday, October 16, 2009

Frame Relay Traffic Shapping Terms

I was reading about traffic shaping & policing and the acronyms in the book they were as clear as mud, but now Ive written them down and understand them they are really really simple. I put a few formulas on for them aswell, but havnt checked them so please correct me if im wrong, oh and I have presumed that you are attempting to traffic shape to the CIR.

Tc – This is a time interval in milliseconds when a Committed Burst (Bc) can get sent. Usually Tc = Bc / CIR

Bc – Committed Burst this is the amount of data in bits which can bet sent every Tc. Usually Bc = CIR / Tc

Be – Excess Burst is the number of bits the Bc can be exceed by if no data has been sent if no data has been sent in previous Tcs. EDIT: As commented by Jeriel Atienza the formula is Be = (Ar – CIR) * Tc/1000

CIR – Committed Information Rate this is the bandwidth of a link or VC in bps which the Service Provider guarantees to provide. Quite often the CIR is lower than the full capabilities of a link which is the main reason why traffic should be shaped & policed. CIR = Bc * Tc

Shaped Rate – This is the rate of the traffic which is being shaped in bps, it normally matches the CIR. Usually CIR = Shaped Rate!

Tuesday, May 12, 2009

Difference between interface service policy(QOS) and inter-zone security policy(ZBF)

The zone-based firewall uses security policy-maps to specify how the flows between zones should be handled based on their traffic classes. The obvious actions that you can use in the security policy are pass, drop and inspect, but there’s also the police action and one of the interesting question is: “why would you need the police action in the security policy if you already have QoS policing”.

The difference between interface service policy and inter-zone security policy is in the traffic aggregation: the interface service policy works on traffic classes entering or leaving a single interface and the inter-zone policy works on aggregate traffic between zones, including the return traffic if you’ve used the inspect command to configure stateful inspection of the traffic class.

For example, you could limit the amount of HTTP traffic between your internal clients and your DMZ segment to prevent the internal users from overloading your public web servers.

Wednesday, January 28, 2009

QOS order of operation

From Cisco's Website I got this important inputs on QOS order of operation:


Inbound
1. QoS Policy Propagation through Border Gateway Protocol (BGP) (QPPB)
2. Input common classification
3. Input ACLs
4. Input marking (class-based marking or Committed Access Rate (CAR))
5. Input policing (through a class-based policer or CAR)
6. IP Security (IPSec)
7. Cisco Express Forwarding (CEF) or Fast Switching

Outbound
1. CEF or Fast Switching
2. Output common classification
3. Output ACLs
4. Output marking
5. Output policing (through a class-based policer or CAR)
6. Queueing (Class-Based Weighted Fair Queueing (CBWFQ) and Low Latency Queueing (LLQ)), and Weighted Random Early Detection (WRED)

Best Regards,
Deepak Arora