Friday, November 02, 2007

NaNoWriMo

When I started the National Novel Writing Month contest I wasn't sure what to expect.  After all, I didn't have a story in mind, no plot had immediately come forth pleading that it needed to be written, and most importantly, I didn't have the foggiest idea if I could do it.

Well, I am a ways into it now and I must say that the pep talks they gave on the site were correct:  the story does have a tendency of writing itself.  With only the opening sentence to work with the story slowly started to evolve and grow.  Characters started coming together, pasts started being revealed and the tone of the story started to come out.  (Now if only I could figure out some way of making these notes count towards my word total.)

It was really strange because, in many ways, that's how I've always done my programming as well.  To be honest, I was never one for getting all of the requirements together before building the application.  I would start with the framework and start building the rest of the application in pieces.  Sometimes this caused me no end of trouble because I had built the framework in such a manner as to cause endless rewrites due to a specific requirement.  I learned, however, and I developed better frameworks.  I learned to "steal" code from the best of the applications and re-use it where necessary.  In essence, I started with the barest of bones and built up the application in stages, much like what is happening with the novel.

Some might call it eXtreme Programming, while others might just lump it in with the generic term Agile Development.  All I can say is that it worked for me.  It does not work for everyone.  Indeed, the vast majority of people cannot do it this way because of the unknowns and the, to be honest, the fear of failure.  I've failed at so many things in my life that failing at writing a computer program, something I love doing, just never crossed my mind.

Managers and team leads need to understand that there is not just one type of developer.  You can't go to the store and pick up a box of Generic Developers (now with Vitamin B12) and have them substitute for your Toasty O Developers,  Supervisors, leaders of people, need to understand that there is a wide range of people and that some people, very few, can be left on their own to build the application.  Indeed, interfering with that development process is sometimes more harmful than letting them run loose.

Does this mean that you have an entire team of highly motivated, highly charged, highly independent developers at your disposal?  No, you don't.  The key is finding those that are and nurturing their growth.  Studies have been done with regard to programmer productivity, and anecdotal evidence abounds with stories of Software Heroes.  Suffice to say that they are relatively rare, so the odds of finding more than one or two on your staff is unlikely.  Which in some ways is a relief.

 

P.S.  For those that want to know, the opening sentence is:

“The screaming didn’t start when the lights went out; it only started when the first body hit the floor.”

Tuesday, October 30, 2007

Don't Worry About Performance!

There was a statement made a number of years ago in an ISWG (Integration Services Working Group) meeting which, to summarize, said "Don't worry about performance, we'll take care of that".  While that is probably going to be the epitaph of at least one person, I think it is time to set the record straight.

Worry about performance.

OK, now that we've gone 180 degrees it's time to put some parameters around this.

  • Don't worry about things over which you have no control.  The speed of a DCOM call is something you have no control over, neither is the time required to create a connection for a connection pool, the time required to retrieve session state, nor the time required to process an HTTP request.
  • Do worry about things over which you do have control.  While you can't do anything about the speed of a DCOM call, you can control, to an extent, the number of DCOM calls that you make.  Less chattiness is better.  While you do not have control over the speed of the resources, you have control over how effectively you use those resources.

The UAT and Production environments into which your application will eventually move has a myriad of interconnected pieces that, together, create the environment in which your application will exist.  While you cannot control those resources you control how you use them.  <Al Gore>  Ineffective use of resources and our unwavering dependency on technologies that produce greenhouse gases is threatening our livelihood and our very planet.</Al Gore>  Ineffective use of resources in any ecosystem is bad, whether that ecosystem is our planet, or our clustered production environment.  Infinite resources are not available and never will be, but effective use can be made of existing technology:

  • Up until the late 1990's the computers on board the Space Shuttle flight deck were Intel 8086 chips.
  • The original Xbox had a 733MHz CPU with 64MB of memory and yet it could (still can) outperform peoples desktops.
  • Mission critical systems have been designed with, when compared to modern technology, primitive means.
  • The first spreadsheet I used, Visicalc, ran on a 1 MHz processor.

All of these examples show that you can make something run well in limited circumstance, but you have to want to.

Identity Theft

I guess I can say that I am now a statistic.  You know, one of those millions of people who have been the victims of identity theft.  Let me tell you the story.

When I got home from work on Monday I noticed that I had supposedly been sending out emails from eBay account at lunch.  Within 15 minutes of the start of the emails I received an "A26 TKO Notice: Restored Account" from eBay UK stating that:

It appears your account was accessed by an unauthorized third party and used to send unsolicited emails to other community members, including email offers to sell items outside of eBay. It does not appear that your account was used to list or bid on any items.

The first thing I tried to do was log into my account.  Well, something either eBay UK did or something the hacker did was change my password.  I tried to enter in my answer to the Secret Question, but that didn't work either as the information on my account had been changed.  Following the various prompts on the eBay site I ended up sending them an email telling them what had happened and what the next step should be.

A couple of hours later I got another email that I apparently sent out eight hours after the first round.  Not content to sit by and wait for the email process to work its way through the system I then started scouring the eBay site for a phone number to call.  You know, that is one of the hardest things I had to do!!!!  I followed all the usual routes and ended up with forms to fill out.  I never did get a phone number, so I had to use their "Live Help" facility.  (My reluctance to go with this approach was due in part to a 45 minute wait on the weekend for "live help" from another company, which never even connected with a human being.)  In the case of eBay, however, the wait was less than two minutes, they told me my position in the queue (started at number 5) and the approximate wait time. 

The person who was on the other end of the chat could have been anyone, anywhere in the world.  The fact of the matter is, they looked at the information on my account, the notes they had sent to me and knew that I needed to talk to the Account Security division.  Within 30 seconds I was "chatting" with someone else who had the power to help.  Two minutes later things were fixed and that included changing the password on my account to a "stronger" password.

Was it brute force hacking of my account and password?  Not if this article is correct.

This particular episode was rather benign in that all that really happened was that some emails got sent and I had to change my password.  It could have been worse.  Much worse.  Think of that the next time you sign up for a web site.  Or, more importantly, think of that the next time you are building an externally facing application.  What are you doing to safeguard the information that you keep on your clients?  What are you doing to protect their safety?  Can you honestly say that you've done your best?

Monday, October 29, 2007

Creative Juices

One of the best things to do, in order to keep the brain alert and creative, is to read different things.  In the same vein, writing is a good therapeutic use of brain cells.  It keeps the neurons working and allows you to be more creative in your job.

To that end, I would like to introduce National Novel Writing Month.  In essence, you are being challenged to create a novel (50,000 words) in less than a month. That's 1500 words per day.  You are considered a "winner" if you actually succeed at getting in 50,000 words.  They don't have to be perfect, you just need to try.

In this case it is definitely the journey which is important, not the final product.  By pushing yourself to reach this goal you are going to be exercising a variety of different areas of your brain.  You will need to be creative to come up with a plot (and subplots), with characters that you empathize with and the words that tie all of this together.

So, how does this help you in the IT field?  I think you would be amazed at how much it will help.  People talk about "thinking outside the box" in order to get something done.  The problem isn't so much thinking outside of the box, it's understanding where the box is in the first place!!!  As you start writing the novel you will be able to see the box that you have created around your novel and this insight, this new vision that you've gained, can help you see the boxes that surround your problems.  Being able to see something is the first step in being able to avoid it or, in this case, think outside of it.

Are you suddenly going to see everything in a new light?  No, but by constantly stretching and pushing your own mind you will see the limitations (the box) that you have put around yourself.

Thursday, October 25, 2007

The Dark Side of Objects

The Dark Side of Objects?  (Luke, I am your father.) 

Sometimes you need reins on developers and designers.  Not because they aren't doing a good job, but because if you don't you may end up in a quagmire of objects that no one can understand. Objects are good, but they can be overdone.  Not everything should be an object and not everything lends itself to being objectified.  Sometimes a developers goes too deep when trying to create objects.

When I was learning about objects I had a great mentor who understood the real world boundaries of objects:  when to use them, how to use them and far to decompose them into additional objects.  Shortly after having "seen the light" with regard to objects I was helping a young man (okay, at my age everyone is young) write an application which was actually the sequel to the data entry application I mentioned in the previous note.  He needed to do some funky calculations so he created his own numeric objects.  Instead of using the built in Integer types he decided that he would create his own Number object.  This number object would have a collection of digits  When any calculations needed to be done he would tell one of the digits the operation to be performed and let that digit tell the other digits what to do.  Well, this gave him a method whereby he could perform any simple numeric operation (+-/*) on a number with a precision of his own choosing.  He spent weeks on perfecting this so that his number scheme could handle integers and floating point numbers of any size.  It was truly a work of art. 

And weeks of wasted time.

What he needed to do was multiple two numbers together or add up a series of numbers.  Nothing ever went beyond two decimal points of precision and no amount was greater than one million.  These are all functions built into the darn language and didn't need to be enhanced or made better.  The developer got carried away with objects and objectified everything when it didn't or in this case, shouldn't have been done.

Knowing when to stop using objects is just as important as knowing when to use objects.

Wednesday, October 24, 2007

Long Running Web Services

OK, the world is moving to web services.  I know that.  You know that.  So what more is there to discuss?  PLENTY!! For instance, how long should it take for a web service to complete?



Well, that's kind of a tricky question.  It basically comes down to "what is the web service doing?"  Some things should come back quickly.  Darn quick, in fact.  For instance, if you ask a web service, "What is the name of the person associated with this identifier?" you should be getting a response back in milliseconds.  If you are asking a web service "What course marks did this student get in high school?" you should be getting a response back in milliseconds.  If you are asking a web service "What are the names of all of the people associated with this school district?" you should be getting a response back in milliseconds.



What?  Getting the names of hundreds, potentially thousands of people, in just milliseconds?  Are you nuts?



Read carefully what I wrote "... getting a response back in milliseconds."  A perfectly valid response is: "Thank you for your request.  It is being processed and the results will be made available at a later date."  Web services should not be long running processes.  If they are long running or have the potential to be long running, then you need to implement some sort of callback mechanism to indicate when processing has finished.  This callback mechanism may be an email, a call to another web service with the results, depositing the results in a queue to be picked up later or even a combination of these methods.  Indeed, there are literally dozens of ways to get the response back to the caller.  What is important to understand is that you do not create a web service that has the potential to run for a long period of time.  Ever.  I'm serious about this one.



Other than the fact that you are probably going to hit the HTTP timeout, COM+ timeout, or cause an excessive number of locks to be held in the database, why other reasons could their be?  Well, imagine from a Denial of Service perspective.  If one call to this web service can generate so much synchronous work, what would 10, 100, or even a 1000 simultaneous calls to this web service due to the underlying server?  Crash it?  Cause it to perform so slowly that everyone gets timed out?  Bring down the back end database?  "But, Don, this is going to be an internal web service only.  What's the problem?"  Depending upon which survey you read, anywhere from 10% to 50% of all attacks on web sites are from insiders.  Imagine a disgruntled employee and what damage he could do to the system with just a little bit of knowledge.



While this topic is ostensibly about web services, we should not create any service (COM+, Web, WCF enabled) that takes a long time to execute.  If you are in the least bit confused about whether something should be synchronous or asynchronous in nature, the odds are it should be asynchronous.  Err on the side of caution.

Friday, October 19, 2007

Hubris

I mentioned a solution architect yesterday and how they needed to be the person with the vision, the person that led the team to final solution.  Well, there is one character trait that a lot of solution architects have that needs to be understood and managed.

Hubris.

Defined as "excessive pride", many solution architects are unwilling to admit they are wrong and will go to any lengths to avoid admitting that they made a mistake.  Sometimes these mistakes can be trivial and sometimes they can be the smallest of decisions that has the biggest of impacts.  Case in point: on a very large project I was working on the solution architect decided that the default order the client had been using for the past 75 years was not right.  So, in the system that we were using he decided that we should change the order of the day, month,and year in the fields that we displayed on the screen.   Yes, that's right, instead of YYMMDD or MMDDYY, he chose a new way of ordering the parts of the date.

This may seem trivial, but we were using a code generator to build the cod and another tool to help create the CICS screens (yes, it was a very old system).  As a result, we had to customize the tool to make it work properly.  OK, it was up to me to make it work properly.  Oh, yay.  Suffice to say that we spent significantly more time on making our dates work out than we would have if we had followed, not just what the client had previously used, but a format that was in use in North America.

No one could convince him he was wrong.  No one at all.  He was convinced that he was right and the rest of the world was wrong.  I know what you're thinking, Don, was he really wrong?  Wasn't this sort of decision part of his job?  I would agree with you, but for those of you old enough to remember, the Y2K scare in the IT field was a big thing.  Imagine a solution architect in the early 90's designing a system where you were unable to enter the century!!!!!  We (okay, me again) had to devise complex schemes that would accurately take a 2 digit year and add the correct century to it.

Solution architects are good, sometimes absolutely necessary, but they are also prone to making very silly mistakes.

Thursday, October 11, 2007

The Good Old Days

Why is it that the good stories start with "When I was younger ..."?

Anyway, when I was younger I was working on project to replace an aging mainframe based system with a new web based application.  The web based system utilized the existing database and a new database to create a whole new system that greatly increased the functionality and usability of the entire product.  One of the response time requirements we had was not with regard to individual screens, nor with regard to specific business processes, but with the length of the transaction in the database.  In order to get as much throughput as possible and minimize the amount of locking the requirement was that the new system operate in much the same manner as the old system and provide an average database transaction length of no more than 0.4 seconds.

400 milliseconds.

That is not a lot of time no matter how you look at it.  This 400 millisecond period was the average length over the course of a business day, but did not include any batch or asynchronous processing that occurrred.  This helped us out considerably because we had a lot of short transactions which lowered the average and a smaller number of longer transactions that raised the average.

Man, did we suck when we went live.  Over 2000 milliseconds for the first week and this did not include any of the deadlocks or timeouts that occurred. It took months, actually 18 of them, before we had things down to not just 400 milliseconds, but an average of just over 300 milliseconds. New hardware on the mainframe helped, but so did the fact that we worked really hard at lowering that average and we understood that anything that was going to take a long time was immediately turned into an asynchronous process or even part of a batch run that night.

The users understood this change in philosophy.  In order to get good online processing for everyone involved there was a need to do things asynchronously or in batch.  Processing something in the middle of the day, when everyone is using the system, is not always necessary.  Even if a short turn around is desired, asynchronous processing can be a valuable alternative.

Funny, but this seems remarkably similar to yesterday's note about web services and performance.  See, everything old is new again.  The problem isn't new, but the technology is.  The solution isn't new either, but the will to implement it might be.

Thursday, September 27, 2007

When not to choose the default

I talked recently about how you should leave the defaults the way they are in many cases because, well, for most circumstances they are probably the best values to use.  Sometimes, however, the default doesn't work that well and you need to understand the reasons why changing the default is a good thing.

Suppose you went to a web site and it asked you to fill in 10 fields.  You then hit submit and it came back and told you that field number 1 is suppose to be numeric, not alphanumeric.  You make the change and then submit.  It then comes back and tells you that filed number 4 is suppose to be less than field number 3.  You keep doing this for a number of changes until you finally say "Forget it" and you leave that site forever.  It's happened to me, so I can honestly speak from experience.

My biggest problem with the process wasn't so much the one error message at a time, but rather the fact that there was a round trip to the server for every error.  I had over half a dozen interactions with the server to fill out a darn form!!!  By default .NET sets the controls you place on an ASP.NET page to process interactions at the server (runat="server").  If you provide complete error checking for each page, then this may be a suitable method of operating.  However, if you only respond to the user one error message at a time, this is sure fired way of getting  someone annoyed with you.  And quickly.

To be brutally honest, some error checking should be done at the client side.  If you have a popular application, or even if it is not that popular, there is still a certain amount of overhead involved in getting IIS to receive the request, process the header information, pass the information along to the application pool and then have the application do what it needs to do in order to tell you that the field is only supposed to contain numbers.  Then the whole thread has to go backward, towards the user, in order to give them the message.  It is faster, more efficient, and less costly from an infrastructure point of view if you let the client take care of this sort of data validation.  Your application will also check when it gets the data from the client (don't ever assume the client is sending you perfect data), but many checks can be performed at the client end, decreasing turn around time for error processing, distributing the work load, and, more importantly, providing a better user experience.

Remember, when you're designing your application think in terms of what provides the best user experience.  Think about your experiences, what you've liked or, more importantly, what you've disliked, and go from there.  The default, while usually good, does not have to remain if there is a good reason to change.

Tuesday, September 25, 2007

Validity vs. Reasonableness

While most of our applications do validity checks on data, not all of them do reasonableness checks.  Let me explain the difference.

Data Validation.  Let us suppose you have a number of fields on the screen:  name, address, birth date, phone number, and spouse's name, spouse's address, spouse's birth date, spouse's phone number and a marriage date.  Data validation would ensure that if there is a birth date, it is a valid date.  So in this case it would check to ensure that all of the date fields are valid.  It would also check to ensure that the phone number follows any one of a number of different standards, but predominantly the fact that it is numeric in nature.  You can also extend data validation to more complex tasks such as determining if the postal code is correct.   In general terms, data validation serves to ensure that a single piece of data is a valid for that data type.

Data Reasonableness.  OK, now that we've gotten the basics out of the way, there are still a number of checks that we can perform.  If there is a marriage date, then the date must be a certain time period after the birth date of both parties.  This is not just a simple "if marriageDate > spouseBirthDate then Happiness()".  We need some additional logic to ensure that even if the data is valid, it also must make sense.  Having data make sense is as important as ensuring that it is valid.

While there are many schools of thought on this, most post secondary training lumps both data validation and data reasonableness together under the "validation" banner.  This, unfortunately, has had the effect, in most cases, of putting data reasonableness checks in the background or has the checks embedded deep within the business logic of the application.  In most cases these checks can be done at the UI level, really quickly and prevent a lot of background processing that clogs the servers.  The other big problem is that some of these reasonableness checks are missed because "I thought the client was going to do that".

Just remember, there is a lot of data checking that needs to be done and reasonableness is just another item in the list.

Monday, September 24, 2007

Temporary (Expires yyyy/mm/dd)

My wife renewed her drivers license recently, as she turned 29, again, on September 24th.  She went into the Registry near our place, filled out the form, paid her money, got her picture taken and was given a "temporary" drivers license to last her until her real license was sent to her.  If the real license takes too long her temporary license is going to expire on her and that could cause no end of grief.

My bus takes a little bit longer to get to work in the morning as there is a "temporary" detour on it's normal route.  The city has dug a hole in the road, the entire width of the road, in order to do some emergency repair on the pipes.  They don't have a lot of time to get this done as it impacts a lot of traffic, a lot of buses (both ETS and school) and a lot of residents.  The only way to get to certain houses is to wind your way through back alleys.  Oh yeah, this is definitely going to be temporary.

My daughter has a temporary spacer in her mouth.  One of her baby teeth, incorrectly filled by an earlier dentist, developed some severe damage and needed to be pulled.  A temporary spacer was inserted so that her teeth would grow in their proper place until the adult tooth comes in.  The spacer is going to be removed either when the adult tooth comes in, or when the baby tooth to which it is attached falls out.  No choice, it is temporary.

We have servers, both physical and virtual, which were set up for temporary purposes.  We now have to face the task of upgrading them from NT 4 to something supported.  (OK, I lied, but you get the point, don't you?)  If things are temporary, give us a date and we will set up the system to self destruct on the day after.  If things aren't temporary, for goodness sake, tell us!!!!!  The "Oh, it was temporary, but the client said ..." story is getting old and, quite frankly, has been done better by other teams. 

Remember temporary means that it goes away.  Pick a date.  Any date.  Please....

Weapons of Mass Destruction

America went to war with Iraq because they wanted to find and destroy the Weapons of Mass Destruction.

In many respects that what I do by looking at the way applications are installed, operate and behave when encountering errors:  I'm looking for weapons of mass destruction.

My first IT job was with a construction company and my crowning achievement was an application that accurately allocated work site costs to various job codes in the accounting system on a daily basis.  It was rather tricky due to the fact that the company I worked for was a multinational company and each country, indeed province/state, had different holidays so it needed to take into account when costs should be allocated based on whether or not the previous day had been a holiday or not, in the location where the construction site was located.  Jobs on which 7x24 construction was occurring had other conditions that needed to be met.  All in all it was a masterpiece of software engineering.

Almost.

You see, I was under a tight timeframe for getting this done, as the VP in charge of construction had told the CEO that it would be in place by July 1st.  I only had 4 more weeks to finish the coding and then implement the application in the main production batch jobs.  Time was tight so I did what every rookie (and many seasoned professionals) do when faced with something that is difficult to compute:  I hard coded the answers.  I hard coded the holidays for all jobs sites for the next 18 months inside the application. 

It was simple to do and saved me a lot of trouble, because, you see, I left the company 2 months later, leaving behind a ticking time bomb in their production systems.  In sixteen months things were going to blow up, all because I took the easy way out instead of doing it properly.

Tick, tick, tick, tick ...

Thursday, September 20, 2007

Groceries and Software

When you buy groceries from a grocery store, one of the things that they do for you is pack your groceries into plastic bags so you can take them home.  (OK, some grocery stores don't do that, but they will charge you a couple of pennies to give you a plastic bag so that you can do it yourself.)  When they pack the plastic bag they have certain, spoken and unspoken, rules.  For instance, you don't put a pound ... err, kilo... of hamburger in the same bag as fruits and veggies unless one of them is wrapped in an additional plastic bag.  You don't put a bag of potato chips at the bottom of the bag and potatoes on top of them and eggs are packed flat.

When creating a deployment for DeCo, think of ZIP files as being the similar to plastic bags from the grocery store.  Here are a few simple rules that you can follow to fill those bags:

  • Create a ZIP file for each item type that you are migrating.  For instance, if you are migrating a web components and COM+ components, create two zip files, one for each.
  • In each ZIP file put all of the pieces necessary to do the work for that type of deployment.  If it is a web zip file, include the MSI, any config files and even the documentation. 

The rules are really simple, but make life much easier for everyone.  For instance, by using ZIP files instead of a list of separate files, you reduce the number of times you need to attach a file to the request and you reduce the amount of space needed on the back end to store the files.  We current have over 9760 files that we are tracking from over 6800 deployments and these files take up over 19,680,000,000 bytes.  While some of those files are compressed, not all of them are. 

By using ZIP files you can help us keep a handle on the storage requirements for DeCo as well as making it easier for the Deployment Analyst or DBA to get all of the files they need for the deployment in one simple package.

Friday, September 14, 2007

Schopenhauer's Law of Entropy

So, just what is Schopenhauer's Law of Entropy?  Simply put, it is this:

If you put a spoonful of sewage in a barrel full of wine, you get sewage

So, what does sewage have to do with programming?  It's not sewage that I'm looking at, but rather the concept behind it.  In IT terms, what Schopenhauer is saying is that no matter how good the overall application, if one part doesn't work the whole application gets tarred with the bad brush. 

It is unfortunate that a single poorly designed, written or executed page can make someone believe that the entire application is poor.  Their perception of the application is what is important, not reality.  Kind of scary, isn't it, when perceptions are more important than reality.  But this is what happens in our business and it is something that we need to understand and do our best to influence.

So what influences this perception?  Well, consider this:  two web applications side by side on your desktop.  You push a button on the left one and you get the ASP.NET error page:  unfriendly, cryptic and somewhat unnerving.  You push a button on the right one and you get an error message in English, that explains there is a problem and that steps are being taken to resolve the issue.  Which one would you perceive to be better written and robust? 

How about another example?  You push a button on the left application and you get an error message that says "Unexpected Error.  Press the OK button".  You push a button on the right application and you get an error message that says "Our search engine is currently experiencing some difficulties and is offline.  Please try again later."  Which one do you perceive to be better?  Which one do you think your business clients will think is better?

It's not just one thing (error message or not) that gives you a feeling of confidence when dealing with an application, it is a multitude of little things.  Making things more personalized helps.  Translating from Geek ("Concurrency error") to English ("Someone else has updated the data before you") helps a lot.  Making it seem that you spent some effort to foolproof the system (i.e. don't make every error number in your application the same error number).

No matter how good the rest of your application, one bad move can create sewage.

Initialize All Variables at Their Points of Declaration

I was reading a book recently called Code Craft - The Practice of Writing Excellent Code and one of the comments struck a particular chord with me as it brought back memories of an upgrade that went horribly wrong.  At least for me.

There was a brief section called "Initialize All Variables at Their Points of Declaration".  Now, this may seem self explanatory and quite normal to some people, but others think that this is rather strange.  "Why would I initialize a variable that I may never use?"  The problem is, that not everyone follows the same coding practices in real life.  Sometimes the compilers help/hurt us in this regard.  Back when I was predominantly working on the mainframe, we were switching from an older version of COBOL to COBOL II.  Ooh.  COBOL.  I can see your eyes glazing over.  Stay with me, there is method to my madness.

The process of conversion was really quite simple.  Recompile.  It wasn't that hard.  However, we discovered a little bit of a problem.  When we did our testing we discovered that we were occasionally getting OC7 (data exception) errors when everything should have been working.  Indeed, running the program multiple times against the same data actually generated different results.  After a lot of head scratching we determined that the problem lay in the fact that the old compiler, by default, initialized variables when they were defined.  COBOL II did not do this by default.  When the application was loaded into memory it would occasionally access memory that had been initialized for some other purpose and the program would work.  Other times, however, it was accessing "garbage" and the program would blow up.  If the original developer had initialized the variables in the fist place we never would have had a problem.

So, we made a small change and everything was perfect.

Almost.  Because of how we were doing the upgrade process I had to baby sit the recompilation of 1900 COBOL programs in 4 different environments (7600 recompiles altogether).  Took almost 48 hours to do it and I got almost no sleep, and all because someone failed to initialize a couple of variables.

Tuesday, September 11, 2007

Antacids for the PM

PMs, can you imagine the look on the faces of those developers when I said that they needed to supply you with things like estimates and time sheets?  I bet some of them were knocked off their feet!

I mean it's not like you were asking them for anything difficult.  I mean, how hard can it be to create an estimate for a project?  It's not like it's rocket science.  Every developer should be able to do it.  I mean, the National Gun Registry hit it's target.  Well, maybe it didn't.  Well, what about the Denver luggage handling system?  Another fiasco?  OK, the FBIs Trilogy Project, now that was ... a disaster?

Everybody can look into the past and come up with failed estimate, but sometimes, the Project Manager does it to themselves.  This may be a shock for some of you, but sometimes Project Managers are under different pressures than you realize.  Back when I was younger I worked for a consulting company.  My job was to come up with the technical work plan and estimates for the projects the local office undertook.  I would then, using historical data and some darn fine guessing, come up with what the effort would be on the technical staff for the project (technical staff meaning project DBA and internal technical support). 

One of my proudest, and saddest, estimating moments came for a relatively simple business project that had a number of interesting technical complexities.  My estimate of the technical cost was 179 days.  When combined with the other parts of the project it turned out that this simple application was going to cost the client a lot of money.  In an effort to reduce the impact tot he client the scope was shuffled, but this did not reduce the technical effort that needed to be spent, so in a classic PM moment the number of days was arbitrarily reduced to 79.  I am proud/sad to say that this is one of the few times in my life where I was dead on with regard to the effort.

Project Managers, if your team says that something is going to take a certain amount of time, ask them questions, make sure they understand both the problem and their own estimate, but don't arbitrarily change it unless you know for a fact it can be done for less.  The odds are they understand what needs to be done better than you and by changing their estimate you are telling them that you don't trust them.  If you still feel the estimate is high, have them sit down with you and walk you through the estimating process they used.  But, if the numbers still add up to something you don't like, take a tums.

Monday, September 10, 2007

Load Balancing Failures

It shouldn't come as a surprise to anyone when I say that our Production environment is load balanced.  I have mentioned this before and I will be mentioning it again in the future.  But, for those who may have missed my previous tirades, let me explain the impact of load balancing on Session State.

One of the features of ASP.NET is to store information specific to a browser session (aka user) into something called Session State.  Session State is kind of like the junk drawer you have at home where you have batteries, twist ties, plastic spoons, stud finders and assorted other "stuff".  Session State allows you store what you need to store in order to keep track of where the user is in the application and what data you need to save on their behalf.  The next time the user accesses the application the session state is automatically loaded and you're ready to rock.

There are a number of places to store session state:  In Process, Out of Process or in SQL Server. 

In Process means that Session State is going to be stored in the Application Pool running the web site.  So, if the application pool recycles, all session state is going to be lost.  A number of projects currently use this method and are in danger of losing Session State, if they use Session State, as we use Application Pool recycling to solve a number of application issues.  In addition, if, for some reason, BigIP sends the user to a different web server to service the request then the Session State is not going to be present, potentially causing a number of application failures to occur

Out of Process is where the session state is hosted by ASP.NET in a different process, potentially on a different machine.  While somewhat safer than storing it in the same Application Pool, a problem arises if this service needs to be reset whereby the Session State is again lost.  Indeed, if the process is hosted on the same server as the web site, moving the request to another part of the load balanced cluster is going to be a problem as Session State will not be available for the request.  If Session State is stored on a separate machine then the biggest problem is that of durability of the data.  Any problems with the service may wipe out all session state for all machines.

Storing Session State in SQL Server is the slowest method, but is by far the safest method for durability and the best method when utilized in a cluster.  Each request for Session State goes out to SQL Server to ensure that the latest and greatest version of Session State for that user is retrieved and used.

In our environment we have asked people to use SQL Server Session State, and yet, by looking trough the web.config files of a number of projects I've noticed that they have their Session State set to In Process.  If Session State is actively being used, this is a recipe for disaster.  I urge each project team to take a quick look at their web.config files and change it to use SQL Server instead of In Process.  Even if you don't currently use Session State, you may in the future and this will prevent you from having a nasty accident.

Thursday, September 06, 2007

In Memoriam

There were a number of comments yesterday from people wondering what the heck I was talking about when I told everyone to take a break.

Recently, Wednesday to be precise, we lost a companion that we had known for more than thirteen years.  When my wife and I first moved into our house we found it to be a large, lonely place.  In order to fill up some of the space we went to the SPCA and picked up a pair of cats, a brother and sister pair.  We named him Spike, because of his spiky hair and we named her Willow.

For thirteen years they were our companions.  Through three kids, hundreds of hair balls conveniently coughed up in the middle of the path to the bathroom, tainted cat food scandals (yes we found a few cans), Spike and Willow were there.  Recently Willow had been having some trouble with her hips.  Old age seemed to setting in quite quickly for her and the doctor recommended a special diet for her kidneys and glucosamine for her joints.  She seemed good, for her, for almost two years, but swiftly went down hill about 10 days ago.  A visit to the vet confirmed the worst:  kidney failure, liver disease, and a host of other problems were manifesting themselves at the same time.  When asked how many months she had, the vet told us "one week".

We made the most of the week with Willow and spent a lot of time petting her and keeping her company, much like she had kept us company for thirteen years.  We soon saw, however, that her time had come to let her go.  The whole family had a good cry on Tuesday night.  Wednesday, a friend of the family helped my wife with the final details and we all had another good cry that night.

I enjoy my work.  I enjoy the people the people I work with, even the ones I yell at a lot.  But I also enjoy other things.  My work does not define who I am, however, just what I do for a living.  Sometimes you need to take a break, step back from the work and look at everything around you.  If you've spent so much time working that you can't remember the last time you truly relaxed, take a break.  If you can't remember the last time you hugged someone or something close to you, take a break. 

Enjoy life.

Take a Break

Today's message is short and simple. Tomorrow we'll explain why.

No matter who you are. No matter what you do. Sometimes you need to stop working and enjoy life.

Take a break from work. Take a break from education. Take a break from stress. Just take a break.

Really Simple Solution

Really Simple Syndication.  RSS.

The concept behind RSS is really quite simple:  users, on their own timetable, download an XML file that contains headlines and/or stories on a particular topic.  For instance, I subscribe to a RSS feed that is created based on the Blog of J.D. Meier.  He is the Project Manager for "security and performance on the patterns & practices team".  His blog contains a variety of interesting topics, but I never actually have to visit his web site to get the latest from him.  I have an RSS reader (in this case Outlook 2007, but there are hundreds of others) that periodically goes out and checks his RSS feed to see if there are any changes.  If there is new stuff it shows up in Outlook.  I subscribe to a variety of blogs and web sites this way:

While RSS feeds are not new, their use in business environments is relatively new.  There are some businesses that have an RSS feed per major application and they post outages, tips, tricks and, most importantly, changes that are or have occurred to the application.  This way their business area is not surprised Tuesday morning when a completely new version of their favorite application shows up.  Larger project teams can create an RSS feed that contains status reports, meeting updates or even updates on when the implementation party is occurring.

RSS feeds can solve a number of issues, but it is not a hammer that can be used on every nail.  It has specific benefits in specific situations.  Like everything else, it is a tool that you have in your toolbox, but it is not necessarily a tool that you need to use.

Friday, August 31, 2007

Apologies

My apologies for the flood of posts today, but I haven't updated the external blog in a long time, whereas the internal email was still being distributed daily. I will try to do better in the future.

Part of the Same Team

"We're all part of the same team, right guys?"


A Project Manager sometimes says this to his team when they've made a decision without consulting him and the decision has some repercussions elsewhere in the project:  money, time, or credibility.


A developer might say this to the management team of the project when the Project Manager or Team Lead has committed to a date that the developer knows is unrealistic, unattainable, or even justifiable.


The business area may say this to the project team when the team seems reluctant to embrace the total vision of the project and seems to be cautious, nervous, or even afraid of the impact.


It doesn't matter the perspective, nor does it matter the person who says it, when this statement is said there is an almost instant "us vs. them" mental image that pops into everyone's head.  Well, maybe not everyone.  Some people, some teams, actually work well together.  They understand the impact of their decisions and, if there are far ranging impacts they discuss them with the required people in advance of agreeing to them.  They understand that even though a request seems simple, they should talk it over with the rest of the team in case something is actually much harder than originally thought.  They understand that being part of a team is a good thing and that teamwork can overcome many obstacles.


Each of us has the ability to shape our team.  Each of us has the ability to help guide the team.  This isn't about being a Project Manager directing the team, it is about people being part of a team and committing to the common goals.


Five people working on the same project is not a team.  Five people, sharing the same vision and goals and working together, is a team. 


 

Orientation day

OK, it's probably pretty obvious that I have been pushing education a lot recently.  Well, today my daughter is attending an orientation day at her new Junior High.  When I think back to when I was her age (yes, the world was black and white back then) going to Junior High was a big change.  Instead of staying in one classroom for most of the day I switched from one room to another and even the people I was with changed throughout the day.


I no longer had the advantage of staying with one teacher a little bit longer and picking up on a concept I missed.  I was no responsible for learning it on my own and, in the event I still couldn't get it, only then was I going to talk to the teacher.  This was a big change in how my world operated up until then and it was really scary.  So, I empathize with my daughter.  I know what she is going to be going through and I will do my best to support her. 


This orientation day my daughter is attending is going to go a long way towards making her feel comfortable in her new school and comfortable with the process.


Now, fast forward ten years.  She's graduated from school and has her degree/diploma and has come to work for your project.  What do you have in place as orientation material?  What do you have that will help her get over the initial fear of a new experience?  What processes are in place to help her become as productive as possible in as short a time as possible?  If you're like most of us, the answer is probably "not much".  We all know the need is there, but filling that need just never seems to be a high priority.


The next time you've got a few minutes, think of my daughter, think of other peoples children, joining your project this year, next year or the year after.  What needs to be in place?  What can you do to help?

Education

Do you ever have a few minutes to kill and you're not sure what to do?  Get certified. 


OK, getting certified in something may take longer than a few minutes, but doing a test is an easy way to tell how close you are to the final goal.  For instance, there is a company called Brainbench that lets you write tests to "certify" yourself in various areas.   While many of these exams do cost money, I prefer looking up the "Free" exams.  Through this route I have taken an exam on Shorthand (I passed, but barely), Internet Security, Writing English, Typing, and others.


I've done these exams for a number of reasons, not the least of which is that I want to test myself to see if I actually know a topic.  I've been talking a lot about Education, recently, and how it is important to keep yourself informed about a topic.  The Brainbench site has a number of FREE exams right now on topics like .NET Framework 2.0, RDMBS concepts, Programming Concepts and Software Testing.  While I am not advocating this particular site, I am advocating education. 


If you are more serious about your education you can try for any one of a number of Microsoft certifications . There are a lot of sites that help you out with studying for these exams, with Transcender being one of the oldest companies in the business.  Or, for those who prefer studying at their own pace with a solid reference, most of the Microsoft exams have associated books.  (Imagine that, they charge for the exam and they charge for the book for studying.  What a racket!!!!) 


It doesn't really matter which route you choose, just go out and learn.

SQL Injection

Security of the data is important to every application.  Ensuring that only properly authenticated users receive access and that only properly authorized users view the data is critical to the success of an application.  Unfortunately, there are many ways to get access to an application and some of them are amazingly simple.  For this note, we're going to talk about "SQL Injection" attacks.


Much like the name implies, a SQL Injection attack is the insertion of SQL code into an existing call in order to compromise security.  Essentially what happens is that the application fails to parse the data coming into the application and allows for people to insert SQL code into an existing SQL call to the database.  For details of how this is done, Steve Friedl of UnixWiz.net has an interesting example.


Is this information hard to come by?  No, it's not.  The link above was actually the top one on the list that Google provided to me.  Detailed, step by step instruction on how to break into a poorly secured web site and the information is so easy to follow that even my daughters can try this out at home.  Many organizations have put standards in place to address this issue.  However, standards are only effective if they are followed and they aren't necessarily going to be followed if the person doing the work doesn't understand the reason why.


Essentially, this comes down to education.  Educate yourself on how to break into your system so that you can prevent others from doing so.  This doesn't mean that you need to be a security specialist, but what it does mean is that you should be conscious of the techniques that people use so that you can stop them from being used against you.  Information is the key.  Let's hope that this key is locking things up instead of opening the lock.

Side Benefits

In a recent note we talked about moving historical records out of the main table into a history table or, depending upon the purpose of the historical records, an audit table.  One of comments that I got back was that had a number of additional benefits:




  1. Easier to write code to retrieve data - no fancy date handling required

  2. Easier to use ad hoc reporting tools - same reason

  3. Better performance due to simplified date handling and smaller table sizes (as only most current record kept)

  4. Can control access to current vs historical data easily by restricting access to the various tables

  5. Easier to archive, as you only need to worry about the history table

(Thanks Rob)


It's easy to miss amongst the glitz and glamour of coming up with solutions that everything we do, every decision we make, has multiple ramifications.  What we may do to "simplify" something may cause severe repercussions in other areas, totally negating the positive benefits.  Sometimes we come across a solution that has both positive and negative impacts, but the positive impacts so far outweigh the negative that there doesn't seem to be a reason not to adopt the new approach.

Coming up with alternatives can be quite difficult, which is where "peer review" comes in really handy.  Grab a friend or two, someone who has done some design work before, and show them your design.  Help them understand the problems and the solutions that you've come up with.  Peer reviews are tremendous tools in that they help to validate approaches and ensure that other possibilities have been considered.  (Don't go overboard on documenting your design until after you've had a peer review, however, as the more time you invest in your solution the less likely you are to consider other options.)

Virtualization Technology

I was reading an article recently about virtualization that actually surprised me.  The Collier County School District in Florida is a very big proponent of virtualization technology.  Their technology plan calls for the replacement of traditional desktops with thin clients.  Users would essentially log into a virtualized desktop located at the District's central computing center.  By loading up blade servers with lots of RAM they are trying to get 30 or more desktops per server.


Wow!  Thirty virtual machines per physical host!  We have not been nearly so aggressive, with our biggest servers handling 15 or 16 virtual machines.  Many of our servers are much smaller and we have a correspondingly smaller number of virtual machines.  Right now we have in excess of 190 virtual machines, some of these being used as desktops, while others are used as servers, both in a Development capacity and a Production capacity. 


With the upcoming release of Windows Server 2008, however, we plan to take even more advantage of virutalization technology.  Comments from Microsoft about the software being able to handle 512 virtual machines per physical machine, notwithstanding, we don't plan on hitting that number any time soon.  What we do plan on doing is implementing features that will allow virtual machines to consume more CPU on the box on which they are hosted, features that will allow us to move a virtual machine from one server to another with no interruption to service, features that will allow us to create new virtual machines in minutes, in some cases in an automated fashion to handle heavier workloads.


Virtualization is a proven technology, just talk to any mainframe guy and he can tell you that multiple "operating systems" are run an IBM mainframe every day.  Great strides are being made in this area everyday and when they are ready to use we will be there.

Error Messages

Error message are vitally important to being able to debug an application that is having troubles.  One thing I should mention, though, is that the error message and subsequent call for action need to make sense.  For instance, the following error messages, or the actions they suggest, just don't make sense or don't help to debug the problem:



  • Keyboard not found.  Press F1 to continue.  (I last saw this on an IBM PS/2 model 55SX.  I paid $6000 for a machine which I felt like throwing out the window.)

  • An unexpected error has occurred.  (I last saw this on a number of different production applications in our own shop.  This doesn't help.  Honest.  Any shred of additional detail would be appreciated.)

  • This is impossible.  (Last seen in one of our production applications.  You know, if I've seen it in an error message, it's obviously not impossible.  BTW, I saw 20 occurrences of this.)

  • Invalid effective end data.  (Too bad there are about a dozen effective dates used at this point in the application.  No idea what date is being used or what table is being accessed.  Quick, call for a DBA!!!)

Sometimes we try to hold our clients hand and we use the excuse "Well, we want to make the error message friendly to the user".  Fine, make it friendly, but you can still had more information.  For instance, on the effective date error if you added what date was incorrect you would not only make it more user friendly, you might actually allow the user to solve the problem themselves!!!  The "unexpected error has occurred" message is sometimes a catchall, but you can still add valuable information. 


No, none of these are perfect solutions, but you need to understand that while you might be covering up the sins of the application to the end user, the support personnel have no data to go on in order to fix the problem.  This prolongs the issue and makes the application actually look worse in the long run.  You might want to consider a two part error message:  first part user friendly, second part techie.  You could add "Report the error to the appropriate support personnel and give them the following data:  blah blah blah".  Give the user both parts, but tell him to pass on the second part.  They will appreciate it, as will I.

Single Point of Failure

Single Point of Failure.


There are probably a lot of really nice definitions our there, but I'd like to use my own.  In my world, a single point of failure is:



... a component, hardware or software based, which when it fails will cause the entire system, or an entire subsystem, to become unavailable to the users ...


So, let's give some examples:



  •  An application that only runs on a single web server has the web server as a single point of failure.

  • An application which uses only a single database server (non-clustered) has the database server as a single point of failure.

  • An application that relies on the Internet, but only has a single connection has their ISP connection as a single point of failure.

While we try to cover many of these different aspects when we design applications and infrastructures, sometimes things still don't work.  For instance, in Production we've got clustered web servers, clustered database server, multiple Ethernet connections, redundant DNS servers, RAID disk storage and dozens of other redundant systems.  Sometimes, though, things just go south really fast and in a really bad way.  Recently we had an air conditioning problem with our server room.  We have redundant units that have multiple air conditioners in each unit.  Through a sad set of circumstances we ended up with only 1 of 4 units working. 


No matter what anyone does, there is no such thing as a full proof system.  There will always be some avenue whereby a single point of failure exists.  The target is to identify those areas and work on putting in redundancy, one step at a time.  It is a long process, but nothing worthwhile is ever accomplished quickly.

DataSets vs. DataReaders

I am stepping into heretical territory here, so you will have to pardon my trepidation.  I am going to discuss something over which wars have been fought, reputations destroyed and live ruined.  Yes, you guessed it, I am going to discuss DataSets vs. DataReaders.


There has been much discussion of this topic behind closed doors and even the occasional directive stating that if you are passing large amounts of data from one tier to another, use a DataSet.  DataSets are indeed convenient mechanisms for transporting around a lot of information that can be stored in a table/row manner.  What happens, though, if you are retrieving a single value?  What if you are going to be retrieving data until a specific event occurs (time or data initiated) and then stop processing?  My contention is that these items may be better suited to a DataReader as opposed to a DataSet.


A DataSet is much lighter weight and is actually the underpinnings upon which the DataSet is built.  When you issue the Fill command to a DataSet it uses a DataReader to retrieve all of the data which it then passes back to you.  if you don't need all of the data, however, you just chewed up a lot of processing cycles, processing memory, and your clients time, retrieving data that you are going to throw away.  If you are in a memory constrained situation or a time constrained situation it may be more appropriate to use a DataReader instead as that will give you more control.  Is it difficult to use?  Heck, no.


So, what is that I am advocating?  Education.  Learn the differences between a DataSet and a DataReader and when each is the most appropriate alternative.  Understand the weaknesses of each, not just the strengths.  Then, only then, make an intelligent, informed decision about the right tool to use. 

History Tables

So, what do you do if you want to have high quality data (i.e. no fake dates for effective end dates) but don't really like to use columns that can contain nulls?  Well, for rows that contain effective dates, have you ever thought about using a history table?


If the vast majority of accesses to the table involve just the current data and not historical data, then a history table may solve your problems.  A history table contains all of the "old" rows and as such it will have an effective start date and an effective end date.  No need for nulls here as you precisely what these dates are.  As for the main table, depending upon the application, it may not even need any effective dates at all!!!!  Need the current address?  Just get it from the Address table.  Need an historical address?  Get it from the Address history table.  Going to be doing this a lot?  Put an index on the date/time fields.  (Sorry about that shameless plug for some other posts of mine.)


Is this effective date nirvana?  No, not really.  There are some applications that make effective use of historical data and for them a history table would only make things more complicated.  In other cases, you aren't really keeping track of history, what the effective start and end dates are being used for is for auditing who made what change on what date.  If what you want is audit information, then create an audit table.  Similar in concept to the history table but designed for auditing.


You see, it's not a sin to take a single table and make it two tables.  Indeed, there are really good reasons why you should.  But, if you aren't sure, talk to your DBA.  They can help you out, if only by asking you questions from a different perspective. That alone is worth the price of a visit.

Null Values

What does a null value in a table actually mean?


Well, technically, a null value means that there is no data for this column.  If the column is to capture a birth date, then a null value would mean that you don't know the birth date.  If the column is about the date of death, then a null value would mean that you don't know the date of death.  It does not mean that the person is alive, just that we don't know the date of their death.


One of the more common problems that developers have is that they make a piece of data, or the absence of the data, mean more than it should.  In the above case, if you need to know if the person is dead, you need an additional field ("Deceased"?) that indicates if the person has shuffled off this mortal coil.  The absence of data in the data of death field cannot, under any circumstances be construed as a field indicating that the person is alive.  What if you were told this person was deceased, but you weren't told when?  What do you do?  Put in a fake date of death?


I have a personal pet peeve in this area.  Within the organization(s) we have a number of tables that have effective dates.  There is a start date and an end date.  What many applications have done is put in "2999-12-31 11:59:59 PM" as the effective end date.  (Historical background: prior to more recent releases of Access, this was the maximum date that Access would allow in a date/time field.)  What this means, to me, is that this record will no longer effective as of that date.  We seem to know this in advance.  Indeed, much of the data that we have seems to expire on this date.  I would not want to be in application support on the day after when all of the data in the organization suddenly expires.


Is this truly the effective end date?  No, it's not.  The effective end date is actually null, but this makes coding for the programmers a little more complicated.  It makes the data cleaner and more accurate, but makes it more difficult to program.


I have a personal preference in this area, as I'm sure you can tell, but I will leave it up to you, the reader, to examine the pros and cons and make up their own mind.  Or, if you'd like, wait until the next Daily Migration Note where a potential solution is revealed.

Windows Registry

The Windows Registry has been a miracle of engineering.  It is a miracle that it hasn't collapsed under it's own weight.


Originally the Windows Registry was the solution to the .ini file.  Instead of storing information in .ini files located all over the hard drive this information could be placed in a central registry and accessed from all applications.  So, what's wrong with this?  Well, like a pendulum that swings from one extreme to another, the concept of the Windows Registry was an extreme.  Yes, there are things that should be placed in the Registry.  Things that are common to one or more applications should place that information in the Registry. Things that aren't probably shouldn't be stored in the Registry.


Like all things, the pendulum is swinging rapidly back the other way.  Using .NET we have application configuration files and web configuration files that are placed in the same directory as the application, pretty much neglecting the Registry.  Is this a good thing?  Well, in many respects, yes.  It gets people thinking about what should be in the registry and what shouldn't be in the registry.


Is it as simple as "Registry=Bad.  Config files=Good"?  No, not really, but it is close to the truth.  Unless you need to put something in a location that multiple applications need to access you should probably use a configuration file.  It is simple.  It is easy to deploy.  It works.

SQL Server Best Practices

Along the lines of our SQL Server set of posts, I came across a really interesting web site:  SQL Server Best Practices.  It contains all sorts of material on best practices with SQL Server 2005.


"But Don, what if we're not using SQL Server 2005?"  Ouch.  Mainstream support for SQL Server 2000 ends on April 8th of 2008.  if you are not currently using SQL Server 2005 then I think your first order of business is getting on to SQL Server 2005.  Don't worry about trying to take advantage of all of the SQL Server 2005 features, just get off of the older software!!!!


OK, now that you've come back from those links and have some awesome ideas on how to take advantage of SQL Server 2005, talk to your DBA.  Please don't automatically assume that everything you read on these pages will be available to you or your project.  Don't assume that all of the whiz bang features have been turned on.  Don't assume that you are familiar enough with the feature to understand the impact it will have in our organization.  Don't assume that just because you read it somewhere and that it made a lot of sense that it actually makes a lot of sense for your application.


Talk to your DBA.  If they don't know about this particular gem that you found on page 764 of a book printed in Greek and found on a Russian web site with the slogan "Punish Microsoft", give them the information and let them research.  As they are more familiar with the base tool than most people they will be able to evaluate your request with a keen eye towards keeping the systems up and running and reducing the effort to do so. 


Once again, talk to your DBA.  And remember, do this before you've written a lot of code as your information may make some radical changes to your database or how you access it.  Proactive interaction with the DBA, an awesome thing to behold.

Talking to your DBA

OK, so last week we talked about the "evils" of SQL Server and how they could be solved through better design and understanding of SQL Server.


One of the questions that I was asked was "how do I get this better understanding"?  Well, there are multiple methods:



  • Reading.  Read books on SQL Server and on effective database design.  Notice how I've separated the two.  Understanding how SQL Server works is just as important as the proper design of the database.

  • Education.  Take a course on database design.  (See if you can find someone who has taken the course before you, as there are some courses that are pure trash and you need to avoid those.

  • DBAs.   Talk to your DBA.  They probably know more than you about SQL Server and how it operates, so it is time to take advantage of their knowledge.

Well the last point, talking to your DBA, may seem obvious, too many people wait until they are in trouble before they talk to the DBA.  That puts a lot of added pressure on everyone and usually results in a "quick fix" as opposed to the best solution.  Talk to your DBA early on in the process, preferably before any code is written, but if that can't be accomplished as soon as possible thereafter.  If there are some fundamental changes that need to be made, you want them done early, not when the Director is breathing down your neck saying "Is it done yet?"


 

SQL Server the Root of All Evil

SQL Server is the root of all evil.  SQL Server causes more problems, with more applications than any other part of an enterprise system.  It is inefficient, slow to respond, and is the focal point of more performance issues than anything else.


OK, now that the myths are out of the way let's get to reality.  SQL Server is indeed at the center of many performance issues, but not because of the SQL Server product itself, but because of the usage of the product.  Database design is a key factor in how well SQL Server, or any database server for that matter, can respond to a query.  What columns actually need to be in your table?  Determining whether you need a sequential  MBUN, a GUID or a ROWID may seem trivial, but it has tremendous impact on performance.  Database design extends to more than just what columns belong in what table, but what indexes should be created and what columns should be used for clustering.  (P.S.  If your cluster index is a GUID, OUCH.  If you don't cluster on your data, OUCH.)


The physical structure, however, is just one part of the overall solution.  Having short, concise and well constructed stored procedures is very important.  Understanding that TempDB should not be used if a TABLE variable works is very important to understand.  Coming up with project standards, at the beginning of the project with regard to expected response time is important.  On a previous project we had a threshold for stored procedures set at 1 second so that if anything took longer than 1 second we looked into the reason why.  Some people significantly lower that number so that anything over 200 milliseconds gets investigated.  If you've set a limit of 1 second and you've told the user that no screen will take more than 4 seconds, then you know that you can call, at most, 4 stored procedures.


If you are experiencing "problems" with the database, don't automatically assume that it is the fault of the database.  If the design, construction and implementation of the database have a flaw you may experience problems, but they are not the fault of the hardware, nor of the database engine itself.

Pet Peeves

We all have pet peeves.  You know, something that really irritates you when other may just shrug it off.  I thought I would list some of the pet peeves of the deployment team:



  • Being asked to remove the old version of ABLearning.xxx.yyy when in the Add/Remove program lists it is called Fred.  This causes no end of trouble.  Can we make the names consistent?  If you are installing ABLearning.xxx.yyy, make sure that the uninstall is called the same thing.
  • Seeing people use a new  deployment package when they migrate the same code to a new environment.  DeCo was set up so that the same deployment package could be used to deploy to UAT and Production.  If the package is named properly it is easy to find and easy to re-use.  (Plus it already has the items attached.  Saves time.)
  • Seeing people use the same deployment package over and over and over and over again.  If you are deploying something new, you need a new deployment package.  Re-using an old deployment package seriously confuses people when they look at the history of other deployments that may have been using the same package.
  • People putting everything into one zip file.  It is so much better for everyone if you match the number of zip files with the number of items.  (Oh, the converse is also true, there is no need to add every single file separately, they can be grouped together in a zip file.)
  • Being told that performance is "slow", but no one can actually tell us what "fast" is.
  • Being asked "when is it going to be fixed", when I don't even know what is broken.
  • People disagreeing with me.

OK, that last one is a joke.  Sort of.

Defaults are Good

Defaults are an interesting thing.  They are usually put in place because, the majority of the time, the default makes sense to use.  This applies to a lot of aspects in life.  The default place for the turn signal indicator in your car is done because, for the most part, this is where people expect it and it is where people can make use of it without taking their hands off the steering wheel.  The default location of Entry/Exit doors to a building or store are set to mimic the roads that we drive on:  enter on the right (just like driving on the right for you out-of-towners).


When we build applications we also have defaults.  For instance, there is a default name for the application configuration file for a .NET application.  DON'T CHANGE IT.  There is a default name for the file that web sites look for if one is not specified.  DON'T CHANGE IT.  There is a default for buttons when they are pressed.  DON'T CHANGE IT.  There is a default behaviour for text boxes.  DON'T CHANGE IT.


OK, perhaps "DON'T CHANGE IT" is a little harsh, but it grabs your attention better than "Don't change it unless there is a reasonable payback on an alternative including, but not limited to: increased productivity, decreased deployment effort, increased user satisfaction, increased performance, etc."  Seriously, though, unless there is a good reason why you want to make changes to the defaults, let them be.  They are defaults because they work for the majority of people, so why don't they work for you?

Doing Data Fixes the Right Way

I was having a rather spirited debate with an old friend the other day about data fixes.  (Since he was buying lunch I thought it only polite that I listen while he talk.  That and I didn't want him to take away the food if I disagreed with him.)  For the most part we agreed on many of the items:



  • when possible data fixes should be tested in UAT prior to Production

  • they should be scheduled much like any other deployment

  • business areas should approve the data fix prior to the data being committed

We disagreed, however, about what constitutes a data fix.  While we both agreed on the general principle of "a data fix is used to correct data that is in an invalid state in a database", I preferred to add one additional word "unexpected".  in my definition the data needs to be in "...an invalid and unexpected state ..."  In my friends world he is used to doing data fixes on a daily basis to correct the data issues that crop up in the inventory system that he is maintaining.  The data fixes are pretty much the same SQL run over and over again, with just the input parameters changing.  When I asked him why he just didn't fix the program he complained that management didn't want him to spend time fixing the code because "... changes were coming ..."


If you expect data fixes as a regular part of your daily operations, then you have a problem with your application.  Fix it!!!!  I know that in some circumstances it isn't easy to fix.  There may be forces outside of the control of your application that cause the data to be invalid.  However, in those cases there is an easier fix than the submission of a data fix every time the business area needs some data changed.  Create a page, a really simple page, that accepts the parameters from the business user and executes a stored procedure to make the changes. 


What does this do?  Well, it eliminates the middle (wo)man, the DBA who needs to create and or run the SQL.  It gives the business area more control over what they want to do when they want to do it.  It can provide a full audit trail of who made the change and when.  And, most importantly, it can be implemented very quickly and will pay for itself very quickly.  The effort around a single data fix, no matter how small, consumes a considerable amount of time.  By letting the user do their own "data fix" the developer/DBA can work on fixing the real problem, not just the symptoms. 


I didn't win the argument, however, as my friend is a consultant and gets paid to do the data fixes, so he actually liked the way it was set up.  I did manage to finish lunch before disagreeing with him, however, so from that perspective I won.


 

Enforcing the Rules

To what extent should "the rules" be enforced?  There are some who advocate complete obedience to the letter of the law.  Others advocate following the spirit of the law, while another group says "as long as it doesn't hurt anyone, does it really matter?"  In some respects it is quite contextual:  you don't follow the letter of the law if it is going to hurt someone, but if no one is around and no one is going to get hurt, do I really have to stop at the stop sign?


The Deployment Team has a number of rules in place with regard to deployments and they were published back in March.  These rules are, we thought, relatively straight forward, but they seem to be either misleading to some people or just ignored.  I though I would re-publish them and ensure that everyone is aware of the rules we follow.  For the most part, we follow the letter of the law with regard to these deployments, mainly because it actually saves us a lot of work and reduces confusion.  If there are any questions about them, please let me know.


Deployment Request Criteria


















Documentation attached

Is there appropriate documentation included with the deployment request so that anyone can deploy the application?  For very simple installs the request can be approved and the deployment team notified that more complete documentation required.  Any deployment that requires the installation of uninstall of an MSI must include documentation.

If there is no documentation the request will be rejected.

Documentation accurate

Is the documentation accurate?  A brief review of the documentation can determine if the document is even pertinent to this migration and if it is accurate.

If the documentation is incorrect the request will be rejected.

Installation files are attached

If DeCo is to be used as an audit tool then the installation files need to be attached to the request and not located on a development server. 

If the required files are not attached the request will be rejected.

Installation files are accurate

If DeCo is to be used as an audit tool then the installation files need to be accurate.  This means that to the best of anyone's knowledge, the files need to be accurate in terms of what they are supposed to do.

If the required files are not accurate the request will be rejected.

Standards are followed

The following standards will be enforced:

- Version numbers

- Database naming

- No Administrator rights to applications

While there are others that we would like to enforce, these are the main items at the moment.

If the standards listed above are not followed the request will be rejected. It is expected that the list of standards to be enforced will grow, but project teams will be notified in advance before the standard is enforced.


There are additional requirements for a Production application









Installation files are from a deployment to UAT

We do not move directly into Production except in the most dire of circumstances.  As a result, all of the files that we are moving into Production need to have gone into the UAT environment first. 

If the files have not previously been deployed to UAT the deployment is rejected, unless the deployment is being made to correct a high priority production incident

Business Approvals made by 8:00 AM on the day of migration

If a Production deployment request does not have business contact approval by 8:00 AM on the day of the migration it will be rescheduled for the following business day, with the exception that nothing will be rescheduled for a Friday afternoon.

The development team is free to reschedule the deployment back to the original day, even if it is Friday, but they will be asked to provide a reason why the migration needs to be done immediately and this reason will be forwarded to Rob Schneider and Dawn Quaife. Friday afternoon deployments to Production will need the approval of Rob Schneider or Dawn Quaife.

If a production deployment request is made for the same day the same process as described above (request for a reason and director notification) will be followed.

Deployments to production are not rejected for this item, but the reasons for late approvals or late creation of requests are made available to Directors for further review.


 

 

 

Documentation

In another life I was the "Master of the Methodology".  OK, what that really meant was that I was the one that knew how to access/install the methodology that we were using and as a result all methodology questions were forwarded to me.  I was the expert, by default.


One of the interesting things that was embedded within the methodology, however, was the concept that the phases of the project and even the deliverables themselves all needed to be defined at the beginning of the project.  While you could use the "standard" project template the vast majority of project managers customized, to some degree, the phases and deliverables that the project was going to create.


While in some circles this would be akin to shooting yourself in the foot, the methodology we used not only made allowances for this, but actively encouraged the modification of the methodology to suit the needs of the business area, the development team and the maintenance area.  There were some documents that were produced that were of specific use to one group only, but there were also many documents that were useful to all parties. There were a few simple rules that were used in determining whether or not a document needed to be produced:



  • What documents are needed by the business area to document their vision?  While there are many different types of documents that could be produced, a short list of really meaningful documentation, created in conjunction with the business area, is what usually made it into the project proposal.

  • What documents are needed by the development team to create the vision shown in the business documents?  If it was not needed to help build the application it was not included in the list of deliverables.

  • What documents are needed to maintain the application?  After being on the maintenance side of the equation for many years, there is a limited subset of documents that are actually useful in maintaining and enhancing an application.

  • What documents are needed by the organization to maintain an appropriate level of oversight on the project?  Status reports, revised project plans, revised timelines, etc., are all required to maintain oversight on a project.

A clear understanding of the word "needing" is required by all individuals.  "Nice to have" is not the same as needing and there needs to be a commonality of understanding amongst all parties:  business, development and maintenance.


The larger the project, the more documentation that is going to be required due to the fact that there are more lines of communication that need to be fully developed and understood.  Conversely, the smaller the project, the fewer the lines of communication and the smaller the amount of documentation required.

Definitions

The English language is an amazing thing.  We have multiple words that mean the same thing and single words that can mean multiple things.  Here are some examples were there can be great confusion over the meaning of a word:


Concurrent users.  To the non-technical user this is the number of users actively using the application.  To the technical person it means the number of people using the application at any one point in time.  The biggest difference is that the more technical  definition specifies an exact period.


Application Downtime.  To the business area the application is down if any portion of the application is unavailable to any portion of the target audience.  To others it means that the entire application is unavailable to the entire audience.


Emergency Change.  This is a situation where there is an immediate need for action or the inaction may cause harm to the organization.  To the more technical  this means an all-nighter and having to convince Don that the deployment request is actually an emergency.


Technical Discussion.  To the business area this is a meeting where the technical people want to get together to discuss something without the business area hearing about it.  To the technical people ... well, I guess everyone agrees here.


A common understanding can go a long ways towards making things run a lot smoother.  For everyone.

Solving the Right Problem

One of the hardest things to do it solve the right problem at the right time. 


When investigating a problem you may end up looking at a wide variety of possible solutions.  Some of these solutions are quick fixes while others require a fair amount of effort to implement.  The question is, which one do you propose?


For a crisis, the quick fix is usually the right choice.  Things need to be resolved quickly and the best solution may not be able to solve the problem fast enough.  As a result the quick fix is usually chosen for Production emergencies and rushed through into Production.  Quick fixes are not meant to be permanent solutions, but in many cases they end up being permanent for a variety of reasons.


In less crisis oriented situations, however, the best solution may actually be the resolution of a deeper, more convoluted problem that is actually the root cause of the issue.  Unfortunately, resolving the root cause of a problem may actually be a problem in and of itself.  There may be significant effort and money that needs to be spent in order to resolve the issue in the manner that it should.  Sometimes the problem is so fundamental to the application that it almost appears that you have to re-write the application to make it work as desired.  If this is the case, is this what you should propose?


As with many things in life, it comes down to a business case:  is the cost of implementing the solution less than the cost of living with the quick fix?  If this were strictly a matter of dollars and cents then the answer would be know right away.  Unfortunately the cost of living with the problem is not easily quantifiable.  How do you measure consumer lack of confidence in terms of cost?  How do you measure consumer satisfaction in terms of cost?  in many cases only the business area affected can even hope to determine the cost.  It is our job to present the facts as we know them, the costs as we know them, and let the business decide the ultimate cost.

Dying Software

Recently I talked about how many of the technologies that we are current coming to an end of their support lifecycle.  Failure to upgrade the technology can cause us problems and here is an example.


We currently use a technology called "iSCSI" to give servers additional storage.  The physical drives are on an iSCSI server and the client essentially maps the space that they are given to one or more drive letters on the local machine.  We have an instance where the space we have allocated is divided into two drive letters:  D: & E:.  The problem arises in the fact that after a restart the second drive (E:) does not always re-appear.  If any application on the server is expecting a drive E: there will be a significant problem.


Microsoft was actually able to recreate the problem on their test machines, but, due to the fact that Windows Server 2000, the operating system that we are using on the client machine, has gone past its Mainstream Support end date, they will not be investing any time in resolving the problem.  They gave us a number of options, but it was pretty much "Good luck and don't bother calling again".


In this case one of the workarounds should fix the problem, but the fact is, if this had been a more serious "production is down everyone come help" type of issue we would have been in serious trouble.


Old technology met new technology and the result was a car wreck.  In order to maintain an operational environment we need to continually update both our hardware and our software.  Being unable to update our software because of dependencies on old versions can cause us some serious trouble.  In this case it was more of an inconvenience than a crisis, but I think we should consider ourselves lucky that our experience was this pleasant. 


If you want to take a look for yourself at what Microsoft products are supported the Lifecycle Information page has a lot of information for you.