1862950 Members
1987 Online
110446 Solutions
New Discussion

Re: A simple checklist

 
steven Burgess_2
Honored Contributor

A simple checklist

How many times have you had a problem where the resolution has been staring you in the face but you couldn't see it for light of day

If you were to have a simple checklist what would reside on that list.

ie

Check the permissions on the directory

Points awarded for good replies

Steve
take your time and think things through
23 REPLIES 23
Sukant Naik
Trusted Contributor

Re: A simple checklist


Hi Steve,

Always maintain a list of phone numbers and fax number of all the engineers / Companies who service your servers.

I paste in the lab and on my desk.

-Sukant
Who dares he wins
Paula J Frazer-Campbell
Honored Contributor

Re: A simple checklist

Hi Steve

The very top of my check list is.

CHECK THAT THE REPORTED FAULT ACTUALLY EXISTS.


Users don't you just love them.

;^)

Paula
If you can spell SysAdmin then you is one - anon
Thayanidhi
Honored Contributor

Re: A simple checklist

Backup the configuration files before editing.
Attitude (not aptitude) determines altitude.
John Carr_2
Honored Contributor

Re: A simple checklist

Hi Steve

could this be when one is day dreaming and not paying attention to the task or simply watching the world cup.

cheers
John.
A. Clay Stephenson
Acclaimed Contributor

Re: A simple checklist

0) What is the user doing differenly; What has been installed/modified/removed since this command was last sucessfully executed?
1) Does the problem actually exist as stated? Is it reproducible?
2) What is the EXACT error message. If known, was was errno or $?
3) Does it occur for only a certain user?
4) Does it occur for only a subset of users?
5) Does it occur at only certain times?
6) If you use someone else's PC/workstation, does the problem go away?
----------------------------------------------
At this point the problem is 'real'.

7) bdf - disk space available; file permissions; mountpoint permissions
8) ulimit
9) Network and hostname resolution
10) Possible maxdsiz, maxssiz, maxtsiz limits reached.
11) Hardware problems

Normally, any network/resource/hardware problems would have already been flagged by IT/O (oops VP/O).



If it ain't broke, I can fix that.
James R. Ferguson
Acclaimed Contributor

Re: A simple checklist

Hi Steve:

Too often it happens that one is stymied because one *assumes* certain "facts".

Assume nothing. Review the obvious and look for the simple solution first. If a all possible, describe aloud the problem to someone else. Vocalizing sometimes tells you the obvious thing you're missing.

Regards!

...JRF...

Mike Hassell
Respected Contributor

Re: A simple checklist

Steven,

I know what you're talking about. As the others have already posted there are some basics when troubleshooting any problem that may arise on your system. Here is what I try to stick to:

1. Does the problem actually exist, or is this what a user "thinks" is occuring. In other words, just verify that there is an actual problem.

2. Look for any error messages that could be associated with the problem at hand, from syslog, application, etc.

3. Can you reproduce the probelm? If not, how often does it occur.

4. Try to isolate the problem to a certain user, environment, etc. Rule things out at this point and take detailed notes on the timing of the problem, any time/date stamps can be very useful.

5. Jog your memory to think of past problems that may be assoicated with this one in anyway.

6. If you have some solid error messages to work with, be sure to check newsgroups for others who may have discussed the problem in detail, which may lead to an answer or make you think of something that will help you move further in resloving the issue at hand. In other words, search:

http://groups.google.com

7. Make detailed notes when resolving problems and keep this notebook handy, for easy review when other problems arise. This notebook will help you in more ways then you would think.

The above may not be a simple checklist, but I feel like it's rather useful when trying to resolve any problem that may arise on your system.

The bottom line is that troubleshooting is not an exact science, so you'll have to find a simple checklist tha suits your style of system administration. :-)

Hope that helps.

-Mike
The network is the computer, yeah I stole it from Sun, so what?
Tim D Fulford
Honored Contributor

Re: A simple checklist

More of a story than a checklist. We recently had a problem with a shared fc60, the alternate link was down on all 3 nodes. We therefore assumed that it was the fault of something common... We tried everything that the machines had in common. It actually turned out not to be the case, one of the servers had a broken fiber cable & somehow the fault propagated to the other two. In the final analysis it is suspected that a combination of old patches + old firmware were more to blame than just the cable.

So my addage is
- even if it is theoretically impossible or improbable, does not mean that it will not turn out to be so.
- Old firmware + new software == more unexpected problems than you can shake a stick at.

Tim
-
Vincent Fleming
Honored Contributor

Re: A simple checklist

I find that using "tusc" to see what a process is doing can be very helpful. Knowing what it's doing usually points to where it's going wrong. You can get tusc from:

http://hpux.cict.fr/hppd/hpux/Sysadmin/tusc-7.0/

When all else fails, have someone else to take a look at it for you - often they'll take a different approach to the problem and find that little thing you've been missing in just seconds. Sometimes you get so stuck down a train of thought, that you don't see the obvious answer, which someone else will see right away.
No matter where you go, there you are.
John Carr_2
Honored Contributor

Re: A simple checklist

Steve

how about :

understand the error
check error logs
check system usage
check for previous occurances on hp forum

John.

PS check one is simply not day dreaming.
Vincent Farrugia
Honored Contributor

Re: A simple checklist

Another story...

Once a B180 could not enter X after booting. We spent an hour tweaking on all sorts of config files but to no avail.

PRoblem was we forgot to attach the mouse to the B180...

Vince
Tape Drives RULE!!!
Vicente Sanchez_3
Respected Contributor

Re: A simple checklist

Steve,

Many "problems" could be resolved if mice, lan cables,
terminal, keybords, etc, where chequed before think about some more "problematic" is happenig.

Many "problems" could be resolved if "users" told us all the steps thay have made to the actual situation.

Regards, Vicente.
Peter Kloetgen
Esteemed Contributor

Re: A simple checklist

Hi Steve,

in my opinion the most important thing is to check *all* changes made on the system, before the error occured.

--> which changes where made on the system?
--> is this the first time the error exists?
--> if not, what was the reason for the error?
--> what was the solution for the error?
--> does the error occur permanently or only in some situations, e.g. when running a special software?
--> when the solution is found, *allways* make a documentation of it!

Allways stay on the bright side of life!

Peter

I'm learning here as well as helping
Nick Wickens
Respected Contributor

Re: A simple checklist

My biggest problem with resolution is forgetting to check log files first and making assumptions that if the sysmtom is the same then the resolution is as well.

My list would therefore be -

(1) How serious is the impact to system operation.
(2) Is the problem specific to one person or system wide.

You now know whether to drop everything or prioritise the problem resolution.

(3) Check system logs.
(4) Check filesystems (ie bdf).
(5) Check hardware (ioscan).
(6) Check performance (measureware/glance/sar)
(7) Check connectivity (ping,lanscan,nslookup)(8) Can problem be reproduced ?
(9) Review recent system changes & backout if necessary.(ie patches etc)
(10) Search ITRC forums.
(11) Raise question on forums.
(12) Call for an engineer.

Of course the number one question whenever dealing with users should always be "Is your monitor turned on". It used to amaze me how often a mainframe user would say "The system is down" when it was just that their monitor was off!.
Hats ? We don't need no stinkin' hats !!
Pete Randall
Outstanding Contributor

Re: A simple checklist

Steve,

My checklist:

1. What's changed?

2. What did the user change?

3. What did I forget that I changed?

Pete

Pete
Scott Van Kalken
Esteemed Contributor

Re: A simple checklist

Checklists don't help with FileNet OSARS... they give the generic error "Storage Library Broken" from software.

Had one once that was cactus... two engineers, and myself decided at 3am when the box wouldn't boot because it was looping on detecting SCSI devices it was one of the OSARS (each library and each drive within the library were seen as seperate SCSI devices). Eventually at 3am after trying everything with multi meters, checking firware revisions on drives etc that it was best to just cable around the faulty drive.

This was after 6 attempted reboots and actually watching the Optical drives for the words "SCSI reset" during the boot sequence of the attached unix box.

Turned out to be a faulty drive spewing crap out onto the bus. Obvious for anyone who hadn't been there for 12 hours trying to get the thing to work.

Scott.
Rita C Workman
Honored Contributor

Re: A simple checklist

On those times when things are a bit bewildering and nothing has worked..I tend to do 3 things ...

One is the same as Jim Ferguson...never assume.

Second I like to use Mr. Holmes approach.....if you can't see what the problem is, than determine what the problem isn't..what remains must be the issue.

And 3rd...I walk away for a quick walk around the block. It never ceases to amaze me how fresh air and a quick break restores the brain cells.

======================
"It is usually the simple things in life that tend to mystify us."
======================

Rgrds,
Rita





harry d brown jr
Honored Contributor

Re: A simple checklist


Steve,

How many times? Well, I strive for only making the mistake once, but I have faulted on some issues more than that.

The clue, at least in my simple mind, is to develop a "knowledge tree", either electronic or paper.

Think in generalities in the questions and in the answers. Don???t build exact questions and exact answers. Say for instance you get an fault code of say 0800, don???t just solve that problem/issue, solve any fault code, and don???t do it just for that machine, do it for all.

let's look at this issue: user calls to say that the server has crashed - because it's not responding to their queries/logins. The number of things that could be "wrong" is in the thousands if not millions. Here's some of the things you could do, but don't forget, you aren't the one that "reported" the "issue":

ping the server, check inetd, check syslogs, gateways, and on and on ...

Now you tell the user that everything looks fine on your end, and then you find through querying the user that they called the wrong number, never had an account on that server but their boss wanted them to run a report, they were attempting to login to the wrong server, and on and on ...

good luck...

live free or die
harry
Live Free or Die
Daimian Woznick
Trusted Contributor

Re: A simple checklist

Problem resolution is probably the best way to learn about your system because you are kind of forced to. However, it is probably the most compicated.

One item that should be underlined a few times in the checklist should be to ensure that the files that were modified during the troubleshooting phase are returned to normal (especially the permissions!).

Another item would be to make sure that you don't get careless with a command because the troubleshooting is probably being performed under a UID of 0.

Always check the logs including any application logs that may be on the system. The mail files may also have useful information in them.

If all else fails a lookup from google could come up with suggestions.

The last item on the checklist should probably be STICK WITH IT or YOU CAN DO IT. Tenacity always helps.
V. V. Ravi Kumar_1
Respected Contributor

Re: A simple checklist

hi,

1. run 'top' utility very frequently, to check out any bottle necks.

2. check /var file system size very often.

3. frequently verigy syslog file for any errors

4. make note of all the installations of patches, softwares and any changes made to the system.

5. take system backup everymonth using make_recovery.

6. include all system files in ur regular backups.

7. get the info about system files modified. put the following entry in crontab to get a mail about such files.

00 01 * * * find /etc /usr /sbin /stand -type f -mtime -1|elm -s "System files
modified yesterday on (machine name)" (machnie name)

8. Prepare a document about permissions on all folders in the system.

9. Don't forget to take the backup of any configuration file before modifications.

10. cleanup /tmp regularly. the following entry in crontab will delete all the files in /tmp which were copied/modified 4 days back and were not accessed after.

50 23 * * * find /tmp -type f -atime +4 -exec rm {} \;

regds
ravi
Never Say No

Re: A simple checklist

Add this one, usually when confronted with a problem, we tend to think a little further. Try to think slower and start with the EXACT error message and work your way from there. If you're using a application that keeps logs, check that one first.

Start with the error and try to connect problems ONE at a time.
Steve Horvath
Frequent Advisor

Re: A simple checklist


My adage is "log in and look."

I know this is an old thread, I found it on a search, and it is so funny.

Story: occured in the late 80s, and demonstrates the 2nd reply from Paula:

The system IIRC, was an Amdahl mainframe running emulated UNIX. User community is research based and highly educated in various fields (though not usually in computing :-)

User calls help desk and reports "Internal CPU chronometer data fault" or something like that in pseudo-esoteric speak. SEVERITY 1!

Operations 2nd tier is called, it's escalated to management. Service call placed for hardware support.

I'm in charge day shift Operations, and was called out of a meeting. I "log in and look," system seems fine to me. I call the user and ask how do you know the "CPU chronmeter" is faulty (trying not to laugh).

User says `date` command is off by 3 hours. I have them echo out time zone variable- it's set to left coast time in the .profile. Fix that all up, with an explanation of offsets. Close out case. Emergency downtime canceled.

Later I became the trainer for support people and lesson #1 was always: Log in and look. Lesson #2- verify problem really exists. Followed with above story :)

Was a common occurance in that kind of shop where Operators were intimidated into willy-nilly actions (like unnecessary reboots) by an "intellectual" user community.

Keith Bevan_1
Trusted Contributor

Re: A simple checklist

Steve,

Top of the list has to be :-

Check the problem does not reside between the keyboard and the chair (ie the user).

Finally, make a list of a few things to kick when someone takes 10 seconds to sort it after you spent hours on the problem.

Seriously though, document the evidence & document your testing. How many times do people go round in circles looking at the same thing.

Keith

You are either part of the solution or part of the problem