Tuesday, April 19, 2011
April fool’s day this year brought with it the Rocky Mountain CMG conference. Best view of any CMG show I have been at.
David Halbig from First Data Corp. (also the guy who organizes the show) kicked off with a surprisingly refreshing presentation for a CMG region. He talked about the cloud in an extremely practical way – as well as laying out the management of performance in the cloud.
David laid out some very compelling examples of issues that can potentially emerge when moving to the cloud. David broke up the monitoring approach into:
1. Continuous Monitoring
2. Specialized – intermittent use monitoring (the fire hydrant)
3. Middleware monitoring (application deepdive)
4. Business Transaction Monitoring – for a cross tier understanding of what is going on
David talked about how many vendors say that they are “BTM capable” – but really do not have the “real thing” – his “sure fire” list of requirements included:
• Horizontal view of aggregate and single transactions across all tiers of interest
• Resource consumption information from each monitored tier – basically saying that a network tap based solution cannot cut it.
• Auto-discovery of transaction path (said that some products needed to be “trained” on transaction paths)
• Capture path to non-monitored tiers
• Continuous operation at volume
• Low transaction path overhead
We (Correlsense) also presented; focusing on how Cloud (public/Private), Agile, SOA and Virtualization create an ever changing environment where the only things that stay the same are the user’s expectation of sub second response times and the transactions that run the business.
Although – as usual – the vast majority of the audience was not anticipating and migration to the cloud in the near future – the sessions were very informative because they dealt with managing performance in a constantly changing environment – something that everyone has to deal with.
Tuesday, April 5, 2011
Improving Conversion Rates
Today, the main tools for optimization are dealing with questions like "where the traffic is coming from" and "what people are doing on my site before they convert". The traditional set of tools available for the web analyst are tools like google analytics that provides information about where the visitors are coming from, funnel analytics tools that provide information about the steps in your acquisition funnel, AB and MVT testing tools to run experiments on the traffic to your site and user recording tools and life-cycle tools that let you view your users flow.
None of the above deals with site performance while we all agree there is a significant weight for this issue on your conversion rate.
Few months ago, one of our clients decided to put his site on CDN. His optimization manager told him that like any other change done to the site, they should run a test and evaluate the ROI of such change. They agreed to run a simple AB test on their Brazilian display traffic (traffic from banners) - 50% will be redirected to the regular site and 50% to the CDN version.
Before starting the test, they went to their IT manager and asked him to check using his monitoring tool how long it takes to load each version. The IT manager did not have a robot/agent in brazil so he bought a virtual server there and installed the agent and came back two weeks after with the following results: it took the original site to load 25 seconds while the new version took 6 seconds. Although it was clear 19 seconds is a huge difference, they decided to run the test after all.
Few days after, they got the results: the conversion rate of the new version was 65% LOWER than the original version... Based on the results, it appeared that the longer the user waited, the better chance he will convert (original version: 25 seconds to load, 6.4% conversion rate, new version: 6 seconds to load, 2.2% conversion rate)...
This is obviously wrong. By looking at the amount of page views of each variation they noticed that the amount of pageviews of the original version was much less than the amount of pageviews of the new version (about 1:3 while it should be 1:1) which lead them to the conclusion that because of the load time, most of the users left the page before it was fully loaded and before the web analytics tool sent the event to the server. The users who did wait 25 seconds for the page were more likely to convert...
Another thing they found was that when they measured the actual time it took for the visitor to load the page, it was much more than 25 seconds (closer to 40 seconds). The IT manager did not include the rendering time or the actual time it took the visitor to load all the files (relying on the network speed determined by the server will never be accurate).
You can assume visitors who are waiting 40 seconds for a page are really interested in its content...
The main things you should take from this story are:
- Page load time and abandonment rate during page load should be part of the basic metrics your analysts are using
- You should have a real site performance analytics tool that will capture 100% of your traffic and not count on robots or other sniffing tools
- You have all kinds of KPIs that provide you with a clear view of all of your SLAs and response time through the funnel except for one of the most important steps - when the visitor arrives to your store.
RUM measures the load time of each and every request for 100% of your traffic, from all around the world, and in the most accurate way. It does not rely on robots and synthetic request or calculates the processing time on the server or the size of the response or guessing the network speed of your visitors. Instead, it uses both client and server agents to calculate the actual time it took from the time the user requested the page till the page was fully rendered. It also tells you how many visitors left before getting the page (and before you got a pageview event to your web analytics tool). RUM has a simple yet powerful user interface that was designed not only for IT people.
Please contact us to schedule a demo or download the free addition of the tool.
Monday, November 8, 2010
Videos: Business Transaction Management by Correlsense
Wednesday, August 25, 2010
More Information on Business Transaction Management
Also Check out this video from CEO Oren Elias about their free Real User Monitoring offer.
Monday, February 8, 2010
Why IT Operations is like an Action TV Series
I like watching the series "24", I can’t really explain why. Every time they nearly get the bad guys something wrong happens, there’s some sort of twist in the plot and they need to start all over again. For example, I’m sure that you are familiar with the following classic scene: The CTU (Counter Terrorism Unit) chopper is following a suspect that is driving a black van. The suspect’s van enters a tunnel, but the van doesn’t leave the tunnel. Instead, a number of different vehicles leave the tunnel at the same time, and the suspect is probably in one of them. By the time they figure out that the black van has been left empty in the tunnel – they have already lost the suspect. They shout “We have lost visual!!!” and are back to looking for the bad guys all over again - then they call Jack Bauer…
IT Operations is just like the CTU; the CTU is responsible for making sure that life goes on without any unpleasant surprises. Similarly, IT Operations needs to do the same in its own space and make sure that the business keeps on running and that business transactions are being executed properly and on time.
When something is about to go wrong, the CTU and IT Operations are expected to prevent it before it affects anyone. So they set up the war room, call everyone in, and start doing their detective work to find the needle in the haystack. If they don’t find it and something goes wrong then the results are significant; either people get hurt (in the CTU's case), or business is impacted.
IT Operations' War Chest
So which tools could IT Operations use to find out that there is a problem, identify the root cause of it and resolve the issue?
For example, IT Operations could use HTTP network appliances that help see every HTTP transaction and measure its response time. These network appliances are just like the CTU's choppers, they do not have adequate visibility into the datacenter. They can indicate that something is wrong with the response time of a transaction, but they cannot show why the response time of the transaction is high and cannot provide the visibility needed for resolution.
IT Operations also uses Event Correlation and Analysis (ECA) tools. ECA tools are like CSI detectives (yes… that’s another one I watch…), and rely on other tools to collect information for them, just like a the CSI detective who collects evidence from a crime scene. ECA tools are just as effective as the products that they rely on to provide them with the data. The issue with ECA tools is that, just like in a crime scene, the thief does not usually leave his ID behind, so all you are left with is just clues, and no accurate data to work with.
Additional tools that IT Ops relys on are; dashboards that monitor server resource consumption, J2EE/.Net tools that are capable of performing drill down diagnostics in application and database layers, synthetic transaction tools and Real User Measurement (RUM) tools. With all of these monitoring tools IT Operations still finds itself in a situation where all lights are green while users are complaining about bad response times. In spite of all of the investment in monitoring tools; the infrastructure that IT Operations is accountable for is still unpredictable. Why?
A Simple Example
Perhaps it’s best to take a look at this classic example: One of our customers had a problem with a wire-transfer transaction. The liability for the problem kept on going back and forth between Operations and Applications, who were pointing fingers at each other as to who was responsible for the issue. “All lights are green” said Operations, “We tested the application and it works just fine” said Applications. Simply put, no existing monitoring tool could point out the problem.
So what was the problem? The answer is simple; it appears that, by design, wire-transfers for over $100K were querying the mainframe nearly 100 times, while other transfers would query it only a couple of times. Same end-user, same application, same transaction, but just a single parameter made the transaction take a whole different path, and made the difference between a 3 second and a 2 minute response time.
What Exactly Are You Monitoring?
Now the question is: why can't existing monitoring tools identify the problem? The reason is simple. Traditional monitoring tools monitor the infrastructure and not the transactions. In a complex heterogeneous infrastructure, there are many tools for monitoring each and every component, but no single spinal cord that is able to show how transactions behave across components. None of the tools are able to deterministically correlate a single request coming into a server with all of the associated requests going out of a server and keep on doing so throughout the Transaction Path. Just like the chopper which could not figure out which of the vehicles coming out of the tunnel contained the suspect who came into the tunnel in the first place.
This situation raises some strategic questions regarding your monitoring approach. How effective is a monitoring framework without that business context? Are you supposed to just to make sure the servers are up and running and applications are responding, or is your real goal is to make sure that the business transactions are being executed as intended and on time?
“In God we trust; all others must bring data”
Applications are tricky, transactions are tricky, and they become even trickier in a complex heterogeneous infrastructure that is composed of multiple platforms, operating systems, application nodes, tiers, databases and where communication between components is in different protocols back and forth for every single click of a button by an end-user.
Only by being able to trace each and every single transaction activation throughout its entire path - 100% of the time, for all transactions, across all components - will you be able to systematically collect necessary granular information in order to get business-contextualized visibility into your datacenter. This kind of visibility is a key factor in being able to identify problems effectively when, or even before they arise.
W. Edwards Deming said; “In God we trust; all others must bring data”. I think he was absolutely right. IT Operations can use choppers, or CSI crime-lab detectives, or Jack Bauers. They all have their roles, but when it comes to fast and effective problem identification as well as many other IT related decision making processes (that’s a whole different article…) real accurate data is required – no partial data, no assumptions.
Business Transaction Management provides you with that data, and by doing so, it provides your IT Organization with visibility and predictability. Wouldn’t it be great if you could go to sleep at night knowing that your infrastructure is reliable? That is, unless you want to play the role of the CTU Director…
The 8th season of 24 will be premiered on January 17th, 2010. 'Till then – why don’t you get yourself a Business Transaction Management solution…
Wednesday, July 15, 2009
IT Reliability through Business Transaction Management
- They call into the help desk to complain – usually only after a number of past events where they were un-happy with the application’s reliability.
- An end user monitoring tool measures bad response times
- Strength – enables the monitoring of the end user’s desktop and can measure response times for fat client based applications
- Weakness – must be installed at each desktop
- Strength – no need for end user installation
- Weakness – Javascript needs to be added to web application code
- Strength – easy installation without code modification
- Weakness – still requires end user installation

Sunday, July 5, 2009
IT Reliability through Business Transaction Management
- By planning for tomorrow’s capacity they ensure that the applications that our co-workers and customers rely on will continue to be reliable throughout increased usage and new changes.
- By managing IT performance they enable the day to day reliability of these applications so that revenue can be generated and growth achieved.
Friday, May 15, 2009
BTM & IT Service Management
From managing incidents to managing changes, availability, security, and compliance, BUSINESS TRANSACTION MANAGEMENT has been leveraged by large IT organizations to augment their ITSM activities.
Let’s look at the impact of BUSINESS TRANSACTION MANAGEMENT to each key ITSM process as well as compliance and auditing activities.
Incident Management – BUSINESS TRANSACTION MANAGEMENT feeds to help desk systems business transaction events and
Problem Management – BUSINESS TRANSACTION MANAGEMENT pinpoints exactly where a problem is in the entire data center topology from a business transaction perspective (instead of wasting hours or days pointing fingers in a war room).
Change and Release Management - BUSINESS TRANSACTION MANAGEMENT assesses in pre-production the impact of change to transaction performance by comparing detailed transaction metrics across builds (providing a more holistic view of performance and more granular drill-down information).
Configuration Management - BUSINESS TRANSACTION MANAGEMENT automatically discovers and models your topology based on real transactions, keeping your CMDB up to date (instead of relying on static modeling).
Availability Management - BUSINESS TRANSACTION MANAGEMENT monitors SLA compliance across all tiers (instead of being limited to specific tiers) –
Capacity Management - BUSINESS TRANSACTION MANAGEMENT monitors transaction volume trends and identifies candidates for consolidation.
Continuity Management - BUSINESS TRANSACTION MANAGEMENT provides business transaction measurements for your disaster recovery testing and failover scenario management.
Security Management - BUSINESS TRANSACTION MANAGEMENT tracks all transactions including the location they are coming from, exactly where they are going and what they are doing – powerful data for both detection of security risks and forensic analysis.
Compliance Management - BUSINESS TRANSACTION MANAGEMENT leverages its full transaction monitoring data to prove regulatory compliance even across virtual servers. And it provides application developers production metrics without access to production (SoD).
Auditing - BUSINESS TRANSACTION MANAGEMENT provides a full transaction audit trail and trending metrics.
Saturday, April 25, 2009
Configuration Management Data Base (CMDB) and BTM
Any organization that has or is planning to implement the ITIL methodology will find great value in a Business Transaction Management (BTM) solution. BTM contributes significantly to the population and utilization of your CMDB, along with the auto creation of a service model relationship within your CMDB.
CMDB Population: Business Transaction Management solutions auto discover your application dependency map by monitoring the real transactions that run through an application's full topology. Additionally, the auto discovered transaction types are then added to the CMDB. Anyone accessing the CMDB that wants to see more specific information about various transactions will be referred directly to the BTM tool's transaction repository.
Enhanced Utilization: BTM solutions enable better utilization of a CMDB by linking true business activities to IT processes, enabling the management of IT from the business perspective. Additionally, BTM solutions connect specific transaction segments and servers – virtual or real – to service degradations enabling rapid resolution.
BTM Enables the Real Time Population of Your CMDB
BTM solutions are designed to operate around the clock in production environments and collect important data on all transactions that flow through all of the IT components in the datacenter. This enables your CMDB to be continuously populated with the critical business processes that are linked to their underlying IT counterparts.
Some examples of Configuration Items (CIs) that a BTM solution can populate a CMDB with are:
› Business services (transaction types)
› Data flows - the full structure of the infrastructure that supports each transaction - down to the methods that are executed - enabling much more granular impact analyses in change management
› The Topology of each transaction type (true service models)
› IP addresses of clients and servers
› All of the possible physical servers that transactions may run through
› All of the possible virtual servers that may be brought up on the fly in response to increasing load
The advantage of defining business relationship CIs with the help of a BTM solution is that those relationships are 100% accurate and are based on the flow of real transaction activations in the production environment, as opposed to other CMDB populating tools which must rely on assumptions, or manual input.
It is now possible to understand which changes affect which transactions – or business services - so that the criticalities of each change can be better understood from a true business perspective.
CMDB and the Help Desk Ticket
A BTM solution will automatically detect slow or hanging business transactions and open an Incident. By doing so, the IT Service Support or Help Desk team can be proactive in dealing with the incident. They can proactively start working on the root cause of the problem and fix it – possibly even before any users feel the negative impact of the degradation.
Tickets that are opened by the help desk are automatically put in the context of the CMDB since the BTM solution links the IT components that are involved in the transaction degradation to that ticket. This not only enables tickets to be sent to the appropriate administrator for resolution, it also enables that administrator to utilize CMDB information to help resolve the issue.
BTM – CMDB Scenarios
A Service degradation
› The BTM solution identifies the degradation of a specific transaction type’s SLA
› An incident is created and a help desk ticket is opened - flagging the server which is causing the latency - for example – the latency is due to a specific application server
› The Help Desk forwards the ticket to the Application Team as part of the problem resolution process
› The Application Team can now look at the change history of that specific application server within the CMDB and identify the cause of the problem
A Change is Being Contemplated
› A change to a specific application in the application server is being contemplated
› The Change Manager identifies all of the real business transactions that are associated with the change
› The Change Manager will be able to anticipate the business impact that the change will have on the business users and will later be able to measure the impact with real user measurements
Prioritization of Incidents
› Two servers fail at the exact same time
› The system administrator knows which server is more critical to the business and can prioritize the recovery - something that is not possible to perform without the true model
Change Verification
› Measuring the impact of a change that has been conducted
› A change to the application has been implemented
› The BTM solution will show the impact of that change by comparing the performance of transactions and transaction segments before and after the change
Roll Back Has Been Performed
› The performance of all transactions in the current time frame is compared to the time frame of the last known consistent state in order to verify the roll back
Comparisons
› The performance of specific transaction segments on two identical servers is compared in order to verify the consistency of performance
› The topologies of two similar applications are compared in order to validate the quality of implementation
A CMDB for the Cloud
Cloud Computing's value proposition to the enterprise is a substantial one - maintaining a CMDB for the cloud introduces new challenges. These challenges are related to the real time scalability of applications in the cloud.
Business Transaction management solutions can at any given moment provide a real time snapshot of the cloud enabling a CMDB to keep up with the constant change.
Additionally, BTM's ability to effectively support a federated environment contributes to the stability and scalability of a CMDB in the cloud.
The CMDB can define a template that is fed in by the Business Transaction Management solution. In this manner the current status of the cloud can be shown – this can only be done with a real time BTM tool.
BTM Brings Value to Your CMDB Investment
Consider the following scenario; an application utilizes twenty different servers within the datacenter, a specific business critical service that the application provides really only utilizes only four of those servers. When building a CMDB with traditional tools it is nearly impossible to crack this scenario since most models are based on theories and assumptions. The only way of achieving this accurate level of granularity with the CMDB – short of going line by line through the application’s code - is by using BTM to help populate the CMDB with real data.
Implementing a CMDB is not a small investment, not to mention maintaining it, utilizing it, keeping it up to date and ensuring its validity. Business Transaction Management solutions enable organizations to better utilize and maintain their CMDB by filling in the gaps that occur in real time. BTM solutions are the only solutions that can link the degradation of a business service to the problematic IT component. This enables the location of the relevant CMDB data needed in order to resolve the problem.
Friday, April 10, 2009
BTM and Real User Measurements (RUM)
A complete BTM solution cannot ignore the “first mile” – the segment from the end-user to the data center. Real User Measurements (RUM) should be part of any complete BTM solution, since the origin of every user related transaction, is at the user.
Enterprises must have the ability to measure the level of service that their users are actually receiving - from their own desktop - and provide fast answers when performance degradation occurs.
Whether response times are degrading, or transaction failures are proliferating, IT staff must provide fast answers.
RUM must address two separate needs:
- Web based applications used by remote users at their home or on a mobile device
- Internal corporate users with desktop or Citrix based applications
For home users, installation on the home user desktop is usually not an option. BTM must provide the technology for measuring latencies from the home user’s browser, without installation on the desktop.
Enterprise users, tend to use numerous applications, some based on fat clients (no browser) installed on their desktop. In such cases a local agent has to be deployed to capture every transaction issued by the user in order to measure its latency and successful completion.
Agent based RUM solutions also provide inventory metrics – how many applications are actually being used, what applications are installed and executed within the user’s operating system, and so forth.