<html>
<head>
<meta content="text/html; charset=windows-1252"
http-equiv="Content-Type">
</head>
<body bgcolor="#FFFFFF" text="#000000">
Hi,<br>
<br>
I disagree, seeing this pattern first hand in resulting large
penalty payment on one case (inability to prove that monitoring
reported up on certain time when it wasn't responsive). If there's a
need to monitor availability, there must be recorded event flow, eg:<br>
<br>
Up 11:00<br>
Up 11:05<br>
Up 11:10<br>
Down 11:15<br>
Down 11:20<br>
Up 11:25<br>
<br>
And why? If we only record:<br>
<br>
Up 11:00 - 11:10<br>
Down 11:15 - 11:20<br>
Up 11:25 - ..<br>
<br>
There's a difference. Did the monitoring system report 11:05 or was
the monitoring system down 11:05? In many cases a need to prove that
system was in certain state in certain times. Especially when there
are SLA disagreements. Even more importantly, what happened between
11:10 - 11:15 and 11:20 - 11:25?<br>
<br>
The SLA behaviour requires that some system noticed something
happening at certain point of time. And if there's issues with
system responsibility say at 11:05, there needs to be a event
showing that the system really did report up.<br>
<br>
- Micke<br>
<br>
<br>
<div class="moz-cite-prefix">On 13.02.2015 14:01, Thomas Heute
wrote:<br>
</div>
<blockquote cite="mid:54DDE780.1050207@redhat.com" type="cite">
<br>
Getting back to availability discussion...
<br>
<br>
To me availability is a set of periods, not so much "time series"
and we should just record change of status (closing the previous
event and opening a new one).
<br>
<br>
- Server is up from 8:00am to 11:30am
<br>
- Server is down from 11:30am to 11:32am
<br>
- Server is unknown from 11:32am to 12:00pm (an agent running
on a machine can tell if a server is up or down, if the agent dies
then we don't know if the server is up or down)
<br>
- Server is in erratic state from 12:00pm to 12:30pm (agent
reports down every few requests)
<br>
<br>
We were discussing the best way to represent availability over
time in a graph, representation in RHQ [1] is very decent IMO, can
be extended with more colors to reflect how often/long the website
was down for each "brick" (if the line represent a year with 52
blocks, 1 block can be more or less red depending on how long it
was done during the week).
<br>
<br>
But thinking of it more, availability graph is not that
interesting by itself IMO and more interesting in the context of
other values.
<br>
I attached a mockup of what I meant, a red area is displayed on
response time graph, that means that the system is down, obviously
there is no response time reported anymore in that period. Earlier
there is an erratic area, seems related to higher response time ;)
Rest of the time the system is just up and running...
<br>
<br>
Additionally I would want to see reports of availability:
<br>
- overall availability over a period of time (a day, a month,
a year...). "99.99% available in the past month"
<br>
- lists of the down periods with start dates and duration for
a particular resource or set of resources (filtering options)
<br>
<br>
Thoughts ?
<br>
<br>
[1]
<a class="moz-txt-link-freetext" href="http://3.bp.blogspot.com/-0MsmG5h5i5E/TfjTMZlvx3I/AAAAAAAAABU/6PKDs0RlzuI/s1600/ProblemManagement-RHQ.png">http://3.bp.blogspot.com/-0MsmG5h5i5E/TfjTMZlvx3I/AAAAAAAAABU/6PKDs0RlzuI/s1600/ProblemManagement-RHQ.png</a><br>
<br>
Thomas
<br>
<br>
<fieldset class="mimeAttachmentHeader"></fieldset>
<br>
<pre wrap="">_______________________________________________
hawkular-dev mailing list
<a class="moz-txt-link-abbreviated" href="mailto:hawkular-dev@lists.jboss.org">hawkular-dev@lists.jboss.org</a>
<a class="moz-txt-link-freetext" href="https://lists.jboss.org/mailman/listinfo/hawkular-dev">https://lists.jboss.org/mailman/listinfo/hawkular-dev</a>
</pre>
</blockquote>
<br>
</body>
</html>