Tuesday, April 21, 2009

Risk assessment and proposal for VMWare

Problem, and risk assessment in our current environment and suggested solutions for them.

Problem – Rack Space is limited, hardware utilization is at 10%, cost per server roll is high which creates a struggle for prioritizing and justifying IT projects, and power consumption is unnecessarily very high.
Solution – All of these issues are resolved with VMWare infrastructure server.
VMWare infrastructure adds 70% hardware utilization which allows for 95% less hardware. This reduces cost in hardware that will pay for the DR solution and reduce cost of environment and decrease space utilization and battery / power consumption.

Problem - Current environment lacks flexibility. Implementation for new servers is time consuming and dependant on “build to order” servers for new projects. Demand for IT projects greatly exceeds IT capacity. IT is redeploying IT resources (which causes further delays).
Solution – Virtualization adds the flexibility by harnessing the full power of our hardware, decreasing server provisioning time, and freedom to test configuration changes and new solutions without commitment of hardware expenses.

Problem - Current Maintenance and testing is done on “test” servers which can add unnecessary hardware expense, its manual, time consuming and difficult to recover in the event of problem.
Solution - VMotion and DRS.
VMotion and DRS enable non-disruptive maintenance
a. enables VM migration without downtime for applications and users.
b. DRS makes it easy to perform server maintenance without downtime for applications and users.
c. Snap shot technology gives us recovery in a few minutes from any change made to a server instead of a few days.
d. Test servers that are virtualized don’t use hardware resources unless they’re turned on. Test servers that have been virtualized don’t need a separate “test” server set aside through the ability to recover from snapshots.

Problem- Complex / hardware dependent environments create unreliability. Out of warranty production servers make us vulnerable to high amounts of downtime due to lack of resources.
- Fact: 1 of 4 organizations had significant disruption in their systems. 24% of those outages were > 24 hours
- Fact: Almost 60% of surveyed companies incurred significant financial trouble as a result.
Current solution- Rebuilding servers from backup, if out of warranty hardware fails we don’t have a plan in place. This creates slow, significant downtime, is hardware, driver, firmware dependent.
Potential solution- Standby server – Expensive, hard to maintain
Current solution- Clustering – Complexity
Proposed Solution - Using virtual hardware coupled with the HA (High Availability) and VMotion components of VMWare. Added benefits of these solutions include;

a. High Availability – automatically restarts VM’s.
b. Hardware independent.
c. Easy to implement and configure.
d. Replication of system state ensures a VM has all it needs to startup.
e. Easier testing.
f. More reliable
G. Achieve company recovery time objectives.
H. Make disaster recovery faster and more reliable through automation.

Problem - Complex OS and software recovery processes creates vulnerability and added downtime when OS and software corruption occur.

Current Solution – Attempt to resolve through diagnostics and as a last resort we’ll do a manual reinstallation and configuration.
Proposed Solution-
Software and OS recovery issue is resolved with VMWare Infrastructure backups.

Non-disruptive to applications and users

Provide Off-host backup using standard backup software

Restoration is simplified and more reliable.

Time to restore or resolve takes minutes instead of hours or days




Tuesday, March 31, 2009

VMware ESX storage: How to get local storage to act as a raw disk for VMs

http://itknowledgeexchange.techtarget.com/virtualization-pro/tag/vmware-esx/

I assume this would be helpful if you want to configure a cluster in VM's... not sure just found this article and figured i'd save it for use later.

Wednesday, March 25, 2009

Questions I have about VMWare / virtualization

1. Q How does P2V work?
Answer - This is a very general question that would require a full post to explain however there are explanation on VMware's site. I believe there is an acredidation you can test for and acquire to show you know the product. I don't know if you'd get this on top of getting certified or if you're certified it's redundant.
a. Q If I want to test or find compatibility what do I need to do?
Answer - Go to the software manufacturer and you should easily find compatibility information with VMWare... A lot of large and small companies that I researched had full articles explaining their compatibility and willingness to support their product running on virtual servers. This year Microsoft even gave in and broadcasted their new decision to support VMWare. http://support.microsoft.com/default.aspx/kb/957006
b. Q What is hardware mapping?
c. If I convert a server to a virtual machine is it possible to keep the original physical machine in
in tact and ready to start back up once I've finished testing the virtual version?
i. this is more of a schematics / planning question, whether it's virtual or not the issue is with
the Mac address and how it correlates or interacts w/ the rest of the network.
2. VMWare virtual center
a. Q How can I supply contingency with virtual center? Meaning... it's currently installed and
operating on one server, what happens to the cluster and everything else if the server is
down for any reason?
Answer
This is now included in VMWare VSphere 4 so I feel the information that I've posted below is now obsolete. The quick answer is VMWare's Virtual Center is very difficult to recover from in the event of a hard disk or server failure and I strongly suggest you use the newer version instead of an old version if this is important to you. It's not that you can't recover, it's that it's terribly difficult and manual vs. automated and simple with the new version. Before the new version came out I even noticed Citrix pointing out this issue as a "single point of failure" (even though you stay up and running in the event of a virtual center server going down). Sorry Citrix :).

All the information I posted below is old and only meant to show archived documentation. I wrote this disclaimer 07-21-09

Here's the most recent info and I'm pasting the information in case this person removes their article at some point. My only concern about what's noted in this paragraph is that I don't see a specification that you can move clusters. I only see you can "steal" hosts. I would prefer an option that ensures HA, DRS, VMotion is working properly and I'll have to test this to see if that works or not.

03-27-09 update - Sure enough when I try to add one of the hosts that's already managed by the other virtual center I get this message

"The host is already being managed by IP Address:
Only one VirtualCenter Server may manage this host. If you succeed in adding the host to this VirtualCenter Server, the host will lose its connection to its original management server. Are you sure you want to continue?"

This is only after creating a datacenter in the new virtual center which isn't the same "datacenter" that's arleady setup so I'm not able to move or "steal" the cluster / settings which means i'd have to manually reconfigure them. This isn't a huge deal I wouldn't think however it is a concern. I am going to see if I can actually figure out how to move the cluster and everything else from one machine to another because I need to understand the importance of an actual cluster made for redundancy or perhaps a simple VM which according to the article below that's okay. If it's a VM then I can actually move it to any other hardware i'd like and not only that the redundancy is somewhat built in because of the server contingency w/ the cluster. It just seems a little strange that I'm having to manage something from within itself.

"I get questions from customer who want to setup some kind of redundancy for Virtual Center Server. Some run Virtual Center Server right in a VM but want a physical standby in case they loose the host, others want a cluster solution. This post hopefully answers a few questions you might have about doing this.
You can install another instance of Virtual Center, no problem. If the current VC server is unavailable just add the hosts to the new VC server. You will get a prompt that they are already managed by a VC but you can click OK and “steal” them. This prevents two VCs from managing one host and causing issues with the database.
You can also setup VC in a cluster. Virtual Center Server (the windows service) can be clustered using industry standard solutions, and only 1 license is required when only one instance is active at any given time. Active / Passive clustered configurations can be installed and configured to point to the same Virtual Center database (but only one instance should be active at any given time).
Active / Passive instances of the Virtual Center Management server will also require the following configuration settings to be equivalent-
Both should point to the same database (same ODBC connection setup)
Both should be set to the same “Virtual Center Server ID” (configured through the File->VC Settings menu).
Both should use the same public/private SSL keys (contained in the “C:\Documents and Settings\All Users\Application Data\VMware\VMware\VirtualCenter\SSL” directory)
If VC WebService is enabled, both should use the same configuration file (located at “C:\DocumentsAndSettings\AllUsers\ApplicationData\VMware\VMwareVirtualCenter\VMA\vmaConfig.xml”) "


Below is the old info that I found however it proved to be terribly difficult w/ regard to detaching and attaching the database however I was using SQL Express version so it's possible this doesn't work so well in that scenario.
Answer - Contingency is only available through third party tools. The linked item is an interesting comparison or bitch session from Citrix people about this "single point of failure". Not sure which third party tools you'd use but I'm assuming windows clustering would suffice. Possibly just having Virtual center as a VM would work? I am guessing there's a way to do that and better ways or worse ways. I'll look further. In the meantime I found how you can move your Virtual Center from one server to another.

Here's how per this post

"



  • Take backup of Server A sql database
  • Stop all VMware services on Server A
  • Detach Virtual Center database on Server A
  • Stop all VMware services on Server B
  • Delete the Virtual Center database on Server B (database was empty and was created for the installation of Virtual Center on the new server.
  • Copy the database files from Server A to Server B
  • Attach the database on Server B
  • On your Virtual Center user account grant them DBO access to the newly attached database
  • Start the VMware services on Server B
  • Launch the VI Client form Server B
  • You will notice that after a few minutes the ESX hosts will show disconnected because they still think they are being managed by the old Virtual Center Server
  • Right-click and remove the ESX hosts from the cluster
  • Add the ESX hosts back to the cluster
  • Adding the ESX hosts back to the cluster does not put the VM's into any Resource Pools (Hosts and Clusters View) or Folders (VM and Templates View). Move VM's back to the correct Resource Pools and Folders
  • On each ESX host ensure that the licensing information looks correct
  • Test vMotion
  • Add templates back to inventory
  • Move SysPrep files from old Virtual Center Server to new Virtual Center Server
  • Test deploying VM from template. This did not work for me. I received the error message "The virtual center server is unable to decrypt passwords stored in the customization specification" I had to export the customizations (did this before I moved the server) edit the XML file in a text editor and search for the phrase "

Thursday, March 12, 2009

Offline antivirus tools

So we got a virus on one of the servers in our DMZ and it's causing so much traffic on our firewall that it's causing it to restart thus killing all access to any other server in the DMZ. This led me to find an offline virus scanner. I used this article to ultimately conclude these two items would fit my needs nicely.

1. Kapersky AVP tool - download here - this one you just install and use.
2. Trend Micro Sysclean - download here - this one you have to extract the most recent pattern file which can be downloaded from here. This also appears to be a good spyware scanner and there is a different pattern file for that which you can download from the same place that I linked for the virus pattern file.
3. Looks like Dr. WebCureIt is another one and a lot of people on that link liked it most. - Download here.

More roadblocks with VMWare

Trying to test migration of VM's from one host to the other I get the error;

1. Unable to migrate from to : The VMotion interface is not configured (or is misconfigured) on the source host (IP).

I've been looking around on how to configure VMotion however I haven't seen anything. I'll keep updating this blog as I find the solution.

2. I noticed on the configuration tab VMotion shows - Not Used

First attempt to resolve was following these instructions which worked for a lot of people however the post doesn't explain exactly where to configure the VMKernal network or VMKernal port and i'm not seeing anything that talks about VMotion in these settings

"Have you properly configured the VMKernel network?Is VMotion enabled on the VMKernel port?The option is somewhat buried in the config dialogs. It's in-> Host configuration / Networking-> Properties of the vSwitch that has the VMKernel port attached to it-> Select VMKernel from the port list and click Edit..."

Second attempt and solutions was to found here

It appears I had to configure the VMKernal Network Configuration which I could do by following page 30-33 on this guide. I got the information about this guide as well as further diagnostic tests that can be run to troubleshoot this error from here.

now that I appear to be able to move VM's I want to note I'm getting this warning;
The above warning didn't interfere with migrating VM's to different hosts.

Migration from to : Reverting to snapshot might generate errors (warnings) on the destination host.

PS> I forgot to mention in my previous article that I had to enable Intel VT and No Execute Memory in order to use 64 Bit Operationg System VM's.

on the HP DL380 G5
F10 - Advanced / Processor options / enable both Intel VT and No Execute

Wednesday, March 11, 2009

Roadblocks encountered while configuring VMWare ESX Server, Connecting it to a SAN and configuring HA

This is the first time I configured ESX Server to connect to a SAN and it's also the first time I managed to get HA and DRS working so I figured i'd add the notes for what I experienced before it got working.

First of all we have 2 DL385 G5 servers w/ FC connection to an MSA 2012FC SAN via 2 SAN Fiber switches. We configured everything so there's load balancing on the fiber connections so there's two fiber ports on each server connecting to the Switches (1 to each) then the SAN also has dual ports for each controller which are both connected to each Switch.

The main topics i'm going to talk about include the following.

1. Getting the servers to talk to the SAN. a. pointing to the SAN inside the VMWare software
b. resolving a path error after pointing to the SAN. c. mention multipathing is something I haven't addressed and still need to look at. A. first of all I used this reference tool heavily while trying to figure out how the SAN to get configured with Virtual Disks, Volumes and Luns so I could add the storage in VMWare. (you need to point to external storage if you are going to use the high availability (HA), DRS, Vmotion etc. )B. The problem I encountered is that I received multiple errors inside VMWare after pointing to the storage. The exact error was SCSI: 4506: Cannot find a path to device Vmhba:0:1:2 in a good state. Trying path vmhba0:1:2 (the SAN ID was here). The other error that was related is the I/O error every time i'd try to browse the datastore I would try to upload a file to the store from my computer and I got the I/O error. I looked everywhere and couldn't find an answer that directly solved my problem however I found through a lot of troubleshooting and going through a process of elimination that the issue was caused by the Host port failures happening on the SAN. These were the errors I saw in the event log on the SAN. You'll see below it was going up and then down and I could see the errors occurring while looking at the host port status on the SAN because it goes from green to red on random ports.

03-10 14:45:02 111A14416 Host link up Chan1: 2 Loop IDs, Fabric
03-10 14:44:59 112A14415 Host link down Chan1
03-10 14:44:47 111B14430 Host link up Chan0: 2 Loop IDs, Fabric 03-10 14:44:47 111A14414 Host link up Chan1: 1 Loop ID
03-10 14:44:47 111A14413 Host link up Chan0: 2 Loop IDs, Fabric

The reason this was happening was simple and due to configuration error on our part. We knew we had to do this however I assumed the engineers that plugged everything in had already done it so I didn't think to look for it but the way to resolve it was by disabling interconnect.

This information the manual that I linked to in this article is what resolved the issue. pages 41-42 Configuring FC Host Port Interconnects
"For a dual-controller FC system in a switch attach configuration, host port interconnects are always disabled.""3. Set Internal Host Port Interconnect to Interconnected (enabled) or Straight-through(disabled).The default is Straight-through.This setting affects all host ports on both controllers."
After making this small change the SAN storage started working incredibly well.

2. Setting up HA and DRS.
a. enable root user to login via SSH. (this is required unless you want to go to the server physically to configure what's in step (b).
b. resolving
Go to the service console on the physical server & login
  • vi /etc/ssh/sshd_config
  • Change the line that says PermitRootLogin from “no” to “yes”
  • do service sshd restart

b. I configured the cluster and added the hosts however I received this error.
"Configuration of host IP address is inconsistent on host (IP was here): address resolved to Host misconfigured. IP address of localhost 127.0.0.1 not found on local interfaces and interfaces. "

This issue is a known issue and it's because some additional configuration of the hosts file needs to be done if you want to use HA / Clustering. Here's how to fix it.

PS... I got this from here.

"Login to your ESX hosts with your favorite secure shell program and look at your /etc/hosts file.
The file should look something like this:# Do not remove the following line, or various
# programs that require network functionality
# will fail.
127.0.0.1 localhost.localdomain localhost
192.168.14.2 myesxserver.foo.org
See the last line with the fully qualified domain name (FQDN) of the ESX server beside the IP address of that server? What you want to do is append to that line the shortname of the host as well. What you end up with looks like this:# Do not remove the following line, or various
# programs that require network functionality
# will fail.
127.0.0.1 localhost.localdomain localhost
192.168.14.2 myesxserver.foo.org myesxserver
The typical way to do this is to insert a tab, then the name you chose for your server, up to (but not including) the first dot. You want to add every ESX host machine that is in your cluster to each other’s hosts file. Not only does this make HA much more robust, it makes DNS lookups redundant, and that’s a good thing. Ask yourself, if my DNS has an outage for just 12 seconds, do I really want all of my HA nodes going into isolation mode?
That’s it! Save your changes and exit.
Why do we need to do this? I’m not sure why it helps with VMotion, but HA needs it. HA you see was not written by the same developers as ESX. HA was developed by Legato, which is owned by EMC, as is VMware. It’s a marriage made in heaven, but the devil’s in the details!"

also...

I used this information to assist me as well. However it's not directly related to this error.

1. Run "hostname -v -f " this will show full details on the host. To correct simply run "hostname server.domain.com"
2. Run "vi /etc/hosts". Press "i" to fix whaterver may be incorrect
3. Run "service network restart" and then "service mgmt-vmware restart"
4. Re-enable HA on the clusterhost

After I did all this I managed to get HA operational and I'm currently working with multiple VM's pulling from multiple hosts.





Wednesday, September 03, 2008

Exchange 2003, Outlook Cache, Offline Address Book

New email addresses don’t show up in Outlook immediately after adding them in Exchange.


With Exchange 2003 and Outlook configured in Cached mode there is a known delay from the Offline Address Book in Exchange (OAB). When a new mailbox is created you will not see it in the OAB for 24 hours because the default update for this is 24 hours.

This is different from Oulook name cache because the cache in outlook reveals addresses previously typed in. If you have an old address in our outlook cache and then try to send an email to the person a user will complain that the old address still shows up. This is because they need to clear their outlook name cache. To clear your Cache and update the OAB you need to take the following steps.

To update the OAB in Exchange - This seems intrusive however the users will not notice any system hang while this happens and you can do it during normal business hours. This can take 2 minutes to a few hours depending on the size of your mailbox database.

1. Open Exchange System Manager (ESM).

2. Expand Recipients and then select Offline Address List.

3. Right click the Default Offline Address List on the right hand side and select "Rebuild".

4. You will get a message that "This operation rebuilds the Offline Address List" and you click Yes to continue.

To update the OAB in Outlook – This is also necessary to defeat the 24 hour delay with automatically updating. It can't be done from a central location for all users.

1. Open outlook

2. Click the down arrow next to send/ receive

3. Click "download address book". Once you've done this you should see the new email address in the OAB.

Outlook name cache

This is also not something that can be centrally managed for all users. All users will need to do this personally or you will have to do it for them.

1. Close Outlook and then open up Explorer.

2. Select Tools | Folder Options and click on the View tab.

3. Select Advanced View and check the boxes next to Show Hidden Files and Folders.

4. Open up the Search applet and search for *.NK2 files. There will be a NK2 file for each Outlook profile on the computer and it will be named profilename.NK2.

5. Rename the file to profilename.bak or delete the file. When you open up Outlook, it will create a new NK2 file and start a new nickname list.

Saturday, April 12, 2008

Windows Time Services

Managing Windows Time Services across a network from a central location.

There are 3 commands you run in order to ensure successful time server synchronization.

These commands can be run from the command line or you can create batch files to include all PC and server names in your network in one command. This article assumes you know how to create batch files and opening the command line. This article also assumes you are working in an Active Directory environment and that you know which server is your PDC emulator.

Preliminary steps;
Make sure Windows Time Service is enabled.

  1. From your local computer. Go to start/run and type services.msc and hit enter.
  2. The MMC pops up and you can right click services and select connect to another computer, next to another computer type the name of the machine you want to see and select ok.
  3. Now you'll notice in the top left that the computer you're looking at is in parenthesis. Scroll down to the bottom and look at windows time. Now enable it and specify automatic as the startup type if this is not already the case. It won't let you if Domain Time Manager is installed and already configured in which case you'll need to disable that service before you can enable this.
  4. Make sure the PDC emulator has its own SNTP server. The Time Server's SNTP server needs to be something like nist1-ny.WiTime.net (list currently at http://tf.nist.gov/service/time-servers.html)

    You can view the SNTP server on any machine by typing

    net time \\(computer that you want to see) /querysntp

Now you can run these commands from your computer, it's not necessary to run them on the local machine. The key is to follow the instructions in parenthesis because they specify the computer you're working on and the computer you're synchronizing from. One thing you cannot do remotely is change whether or not the computer is selected to automatically update with daylight savings time. That is impossible.

Now these three commands below will help you set the appropriate time services for any remote computers on your network (including the time server).

net time \\(computer that you want to set) /setsntp:(computer that is your time server)

w32tm /resync /computer:(computer that you want to set)

w32tm /config /computer:(computer you want to set)/manualpeerlist:(first time server you want to use),0x1 (second time server that you want to use),0x2 /syncfromflags:MANUAL

The third command includes a way to specify more than one server so you'll have redundancy in case a server goes out. Simple rule of thumb here is to designate a server that's closest to you and that means designating a server on your local network! When you're designating the time server’s time server then specify the one in New York if you're in New York.

Friday, March 14, 2008

the instruction at "0x01b03397" refrenced memory at "0x00000000" the memory cant be "written"

This error kept happening when I would restart the computer and the app it was happening on was explorer.exe.

This resolution I found here

http://www.geekstogo.com/forum/memory-could-not-written-t4911.html

I just chopped a few of the answers together and pasted them below. These were the exact steps I took to resolve them.

Go to Start and then RUN.type cmd (This will bring up the dos window.type cd.. until you get to the c:/ root directorythen type attrib -r -h -s boot.ini follow that by typing edit boot ini

multi(0)disk(0)rdisk(0)partition(1)\WINDOWS="Microsoft Windows XP Professional" /noexecute=alwaysoff /fastdetect

back at the DOS prompt type attrib +r +h +s boot.ini to re enforce the attributesback at the DOS prompt type attrib +r +h +s boot.ini to re enforce the attributes

Wednesday, March 12, 2008

Windows Deployment Services

Here's a good link to how to setup and use WDS.

http://sociallybeta.blogspot.com/2007/03/wds-rocks-complete-guide-to-using-wds.html

Excerpt regarding sysprep

"Create a folder in the root of the C: drive called sysprep
> Put the Windows XP Cd in the drive and navigate to the Support folder
Extract "Deploy" to the sysprep folder on the root of C; that you just made
> Go to the sysprep folder and double click "setupmgr.exe" to create the answer file needed for an automatic installation on Windows."

Friday, September 21, 2007

Configuring Permissions for a "home" directory referenced in Active Directory Users & Computers on the Profile tab.

1. Specify Read access for the domain users group the the root drive (commonly C:\)
2. Create a group for the people that will have a home directory and add those users to it.
3. Create a share for the "home" directory anywhere you'd like.
4. Give the group deny list permissions on the security tab but for the share permissions give the same group change and read permissions.
5. Specify the home directory in AD users & Computers specify (properties of user, profile tab) it as \\server name\share name\%username%

the fifth step will automatically create the folder and give that person exclusive permissions.

Wednesday, September 12, 2007

Locking down windows 2003 terminal server

http://www.microsoft.com/windowsserver2003/techinfo/overview/lockdown.mspx

That's where you can download a basic guide to lock everything down. Some things to remember.

You can lock down the server in 3 ways.
a. without active directory and via local group policy which locks everybody out even the administrator.
b. with active directory & loop back processing enabled then configuring user settings so the same user name can log in to a locked down terminal server without interfering with other group policy permissions. (you do this and the c. option with the terminal server inside of an organizational unit). With this option be sure to go into the permissions of the GPO and select deny for the domain administrator or whoever else you'd like to have regular access to the terminal server.
c. with active directory and loop back processing disabled. you should only do this if your users accessing the server are inside of the locked down GPO and they don't need to access any other node on the network with anything other than these permissions. Some people setup different user accounts in this scenario and you would do that if the person needs regular access to other nodes on the network while being locked down on the TS.

This is the only 3 that I'm aware of that you could possibly need. Hope it helps.

One extra setting I noticed for disabling IE access for the users helps because you can still access IE regardless of whether or not you removed the IE icon from the desktop.

Here are the instructions for that


Disabled Internet access for users


User Configuration \Windows Settings\Internet explorer maintenace->connection->proxy settings
click enable proxy, make sure enable for all protocols is on, the set the server to any non-existant local server (the computer name noserver works fine)
apply the policy

This will just make it so every attempt for a user to access a web page will time out.


Also, you may want to goto the gpo for:
user configuration\Administrative templates->Windows compnents->Internet explorer and enable to gpo for 'disable changing proxy settings'

Friday, August 31, 2007

Slow Remote Desktop Session in Vista?

this was taken from
http://www.onegeek.ca/index2.php?option=com_content&do_pdf=1&id=35

because I didn't want to lose itThis was originally posted by Tom Keating on another site.





Here's the original post link, didn't want to lose this baby so I copied it to my site too.





http://blog.tmcnet.com/blog/tom-keating/microsoft/remote-desktop-slow-problem-solved.asp





If you're having issues with an RDP session just being slow, as in delayed in transmitting the desktop or mouse clicks,
and other communications to a perfectly good server, check this out.





Looks like is a QoS in Vista policy.



Remote Desktop slow problem solved



April 19, 2007



Remote Desktop 6.0, the latest version of Microsoft Remote Desktop client, which comes pre-installed with Vista was
slower than molasses when I tried connecting to some Windows 2003 servers. In particular, I was trying to manage a
Windows 2003 R2 64-bit Server running Exchange 2007 with 4GB of RAM and a fast 1.83Ghz dual-core processor. I'd
click on something and wait and wait for my click to register. Moving a Window would also be painfully slow. It didn't
seem related to network connectivity since the screen redraw was fairly fast, but it just took a long time for the server to
respond to keystrokes, mouse-clicks, etc. It had all the earmarks of a server's CPU being overwhelmed.




But surely, this brand-spankin' new server will all this horsepower couldn't possibly be overloaded unless it had spyware
or a virus. That wasn't likely either since I'm pretty diligent about protecting my servers. I logged on locally to the server
and the server's performance was normal. Thus, only when I used Remote Desktop was it slow. Further, when I tried
Remote Desktop from a Windows XP Professional PC, the server was also fast. It was only when I used Remote
Desktop from my brand new Windows Vista Ultimate Edition PC that the performance was terrible. It was very odd,
because from my Vista PC I could connect to many other machines with no problems. I was aware that Vista comes with
a new RDP (Remote Desktop Protocol) client called Remote Desktop 6.0, which has more security and networking
features, so perhaps there was some sort of network security conflict.








After doing some more research I discovered that Remote Desktop 6.0 leverages a new feature called auto-tuning for the
TCP/IP receive window that could be causing the trouble. What is auto-tuning for the TCP/IP receive window? Well, the
new Microsoft TCP/IP stack supports Receive Window Auto-Tuning. Receive Window Auto-Tuning continually
determines the optimal receive window size by measuring the bandwidth-delay product and the application retrieve rate,
and adjusts the maximum receive window size based on changing network conditions.





One Geek
http://www.onegeek.ca Powered by Joomla! Generated: 31 August, 2007, 14:26
In Vista, Receive Window Auto-Tuning enables TCP window scaling by default, allowing up to a 16 MB window size. As
the data flows over the connection, the TCP/IP stack monitors the connection, measures the current bandwidth-delay
product for the connection and the application receive rate, and adjusts the receive window size to optimize throughput.
The new TCP/IP stack no longer uses the TCPWindowSize registry values which many third-party utilities used to
"tweak".




Receive Window Auto-Tuning has a number of benefits. It automatically determines the optimal receive window size on a
per-connection basis. In Windows XP, the TCPWindowSize registry value applies to all connections. Applications no
longer need to specify TCP window sizes through Windows Sockets options. And IT administrators no longer need to
manually configure a TCP receive window size for specific computers.




According to Microsoft, with Receive Window Auto-Tuning, a Windows Vista-based TCP peer will typically advertise
much larger receive window sizes than a Windows XP-based TCP peer. This allows the other TCP peer to fill the pipe to
the Windows Vista-based TCP peer by sending more TCP data segments without having to wait for an ACK (subject to
TCP congestion control). For typical client-based networking traffic such as Web pages or e-mail, the Web server or email
server will be able to send more TCP data more quickly to the client computer, resulting in an overall increase in
network performance. The higher the BDP and application retrieve rate for the connection, the better the performance
increase.




The impact on the network is that a stream of TCP data packets that would normally be sent out at a lower, measured
pace, are sent much faster resulting in a larger spike of network utilization during the data transfer. For Windows XP and
Windows Vista-based computers performing the same data transfer over a long, fat pipe, the same amount of data is
transferred. However, the data transfer for the Windows Vista-based client computer is faster due to the larger receive
window size and the server's ability to fill the pipe from the server to the client.




With better throughput between TCP peers, the utilization of network bandwidth increases during data transfer. If all the
applications are optimized to receive TCP data, then the overall utilization of the network can increase substantially,
making the use of Quality of Service (QoS) more important on networks that are operating at or near capacity. Obviously,
this feature is good for ensuring better Voice over IP quality as well.




In any event, I discovered that Vista's Receive Window Auto-Tuning could have issues on some networks. I really didn't
want to disable Receive Window Auto-Tuning due to it's QoS, bandwidth speed/throughput, and VoIP quality benefits,
but I had no choice. I use Remote Desktop all the time to manage 30+ servers. After disabling Receive Window Auto-
Tuning, the "slowness" problem with mouse-clicks, keystrokes, and screen redraws went away. Problem solved! Woohoo!




Here is what you need to do if you have the same issue:




- Run a command prompt (cmd.exe) as an Administrator


- Type: netsh interface tcp set global autotuninglevel=disabled




If you want to to re-enable it:


- Type: netsh interface tcp set global autotuninglevel=normal




In some cases you may need to use this command in addition to the above, but I didn't have to:
One Geek
http://www.onegeek.ca Powered by Joomla! Generated: 31 August, 2007, 14:26


- Type: netsh interface tcp set global rss=disabled




Now, because Receive Window Auto-Tuning increases network utilization of high-BDP transmission paths, the use of
Quality of Service (QoS) or application send rate throttling is important for networks that are operating at or near
capacity. So I'd like to get this feature working, which will require some network topology examination. I did read that
Windows Vista supports Group Policy-based QoS settings that allow you to define throttling rates for sent traffic on an IP
address or TCP port basis. So perhaps I can just disable auto-tuning for the RDP port 3389 and leave it on for all other
ports.




I'm headed over to Microsoft's site which has some excellent resources on policy-based QoS. From my initial research it
looks like you can configure some pretty nifty QoS policies. For example, you can specify a QoS policy with a DSCP
value of 46 for a VoIP application, allowing routers to place those packets in a low-latency queue, or you can use a QoS
policy to throttle a set of servers' outbound traffic to 512 KBps when sending from TCP port 443 (HTTPS port). In theory,
I can set Remote Desktop to have "top" priority and give it all the bandwidth it needs. Heck, maybe I'll set just my IP
address and my Remote Desktop port to have top priority on our network. To hell with the rest of my fellow co-workers!
They don't need no stinkin' bandwidth. It's mine! All mine!

About Me